{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-human-pose-models-from-synthesized","title":"Learning Human Pose Models from Synthesized Data for Robust RGB-D Action Recognition","arxiv_id":"1707.00823","date":"2017-07-04","proceeding":null,"authors":["Jian Liu","Naveed Akhtar","Ajmal Mian"],"abstract":"We propose Human Pose Models that represent RGB and depth images of human\nposes independent of clothing textures, backgrounds, lighting conditions, body\nshapes and camera viewpoints. Learning such universal models requires training\nimages where all factors are varied for every human pose. Capturing such data\nis prohibitively expensive. Therefore, we develop a framework for synthesizing\nthe training data. First, we learn representative human poses from a large\ncorpus of real motion captured human skeleton data. Next, we fit synthetic 3D\nhumans with different body shapes to each pose and render each from 180 camera\nviewpoints while randomly varying the clothing textures, background and\nlighting. Generative Adversarial Networks are employed to minimize the gap\nbetween synthetic and real image distributions. CNN models are then learned\nthat transfer human poses to a shared high-level invariant space. The learned\nCNN models are then used as invariant feature extractors from real RGB and\ndepth frames of human action videos and the temporal variations are modelled by\nFourier Temporal Pyramid. Finally, linear SVM is used for classification.\nExperiments on three benchmark cross-view human action datasets show that our\nalgorithm outperforms existing methods by significant margins for RGB only and\nRGB-D action recognition.","url_abs":"http://arxiv.org/abs/1707.00823v2","url_pdf":"http://arxiv.org/pdf/1707.00823v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"skeleton-based-action-recognition","task_name":"Skeleton Based Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[{"method_slug":"svm","method_name":"SVM"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/skeleton-based-action-recognition-on-ntu-rgbd","task":"Skeleton Based Action Recognition","dataset":"NTU RGB+D","model":"HPM_RGB+HPM_3D+Traj","rank_in_archive_order":110,"of":135,"metrics":{"Accuracy (CS)":"80.9","Accuracy (CV)":"86.1"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}