{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/skeleton-aided-articulated-motion-generation","title":"Skeleton-aided Articulated Motion Generation","arxiv_id":"1707.01058","date":"2017-07-04","proceeding":null,"authors":["Yichao Yan","Jingwei Xu","Bingbing Ni","Xiaokang Yang"],"abstract":"This work make the first attempt to generate articulated human motion\nsequence from a single image. On the one hand, we utilize paired inputs\nincluding human skeleton information as motion embedding and a single human\nimage as appearance reference, to generate novel motion frames, based on the\nconditional GAN infrastructure. On the other hand, a triplet loss is employed\nto pursue appearance-smoothness between consecutive frames. As the proposed\nframework is capable of jointly exploiting the image appearance space and\narticulated/kinematic motion space, it generates realistic articulated motion\nsequence, in contrast to most previous video generation methods which yield\nblurred motion effects. We test our model on two human action datasets\nincluding KTH and Human3.6M, and the proposed framework generates very\npromising results on both datasets.","url_abs":"http://arxiv.org/abs/1707.01058v2","url_pdf":"http://arxiv.org/pdf/1707.01058v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"gesture-to-gesture-translation","task_name":"Gesture-to-Gesture Translation"},{"task_slug":"motion-generation","task_name":"Motion Generation"},{"task_slug":null,"task_name":"Triplet"},{"task_slug":"video-generation","task_name":"Video Generation"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/gesture-to-gesture-translation-on-ntu-hand","task":"Gesture-to-Gesture Translation","dataset":"NTU Hand Digit","model":"SAMG","rank_in_archive_order":2,"of":6,"metrics":{"AMT":"2.6","IS":"2.4919","PSNR":"28.0185"},"uses_additional_data":false},{"leaderboard":"/sota/gesture-to-gesture-translation-on-senz3d","task":"Gesture-to-Gesture Translation","dataset":"Senz3D","model":"SAMG","rank_in_archive_order":4,"of":6,"metrics":{"AMT":"2.3","IS":"3.3285","PSNR":"26.9545"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}