{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-human-pose-estimation-features-with","title":"Learning Human Pose Estimation Features with Convolutional Networks","arxiv_id":"1312.7302","date":"2013-12-27","proceeding":null,"authors":["Arjun Jain","Jonathan Tompson","Mykhaylo Andriluka","Graham W. Taylor","Christoph Bregler"],"abstract":"This paper introduces a new architecture for human pose estimation using a\nmulti- layer convolutional network architecture and a modified learning\ntechnique that learns low-level features and higher-level weak spatial models.\nUnconstrained human pose estimation is one of the hardest problems in computer\nvision, and our new architecture and learning schema shows significant\nimprovement over the current state-of-the-art results. The main contribution of\nthis paper is showing, for the first time, that a specific variation of deep\nlearning is able to outperform all existing traditional architectures on this\ntask. The paper also discusses several lessons learned while researching\nalternatives, most notably, that it is possible to learn strong low-level\nfeature detectors on features that might even just cover a few pixels in the\nimage. Higher-level spatial models improve somewhat the overall result, but to\na much lesser extent then expected. Many researchers previously argued that the\nkinematic structure and top-down information is crucial for this domain, but\nwith our purely bottom up, and weak spatial model, we could improve other more\ncomplicated architectures that currently produce the best results. This mirrors\nwhat many other researchers, like those in the speech recognition, object\nrecognition, and other domains have experienced.","url_abs":"http://arxiv.org/abs/1312.7302v6","url_pdf":"http://arxiv.org/pdf/1312.7302v6.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-human-pose-estimation-features-with","repo_url":"https://github.com/max-andr/joint-cnn-mrf","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"object-recognition","task_name":"Object Recognition"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1312.7302","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}