{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-to-estimate-pose-by-watching-videos","title":"Learning to Estimate Pose by Watching Videos","arxiv_id":"1704.04081","date":"2017-04-13","proceeding":null,"authors":["Prabuddha Chakraborty","Vinay P. Namboodiri"],"abstract":"In this paper we propose a technique for obtaining coarse pose estimation of\nhumans in an image that does not require any manual supervision. While a\ngeneral unsupervised technique would fail to estimate human pose, we suggest\nthat sufficient information about coarse pose can be obtained by observing\nhuman motion in multiple frames. Specifically, we consider obtaining surrogate\nsupervision through videos as a means for obtaining motion based grouping cues.\nWe supplement the method using a basic object detector that detects persons.\nWith just these components we obtain a rough estimate of the human pose.\n  With these samples for training, we train a fully convolutional neural\nnetwork (FCNN)[20] to obtain accurate dense blob based pose estimation. We show\nthat the results obtained are close to the ground-truth and to the results\nobtained using a fully supervised convolutional pose estimation method [31] as\nevaluated on a challenging dataset [15]. This is further validated by\nevaluating the obtained poses using a pose based action recognition method [5].\nIn this setting we outperform the results as obtained using the baseline method\nthat uses a fully supervised pose estimation algorithm and is competitive with\na new baseline created using convolutional pose estimation with full\nsupervision.","url_abs":"http://arxiv.org/abs/1704.04081v1","url_pdf":"http://arxiv.org/pdf/1704.04081v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-to-estimate-pose-by-watching-videos","repo_url":"https://github.com/prabuddha1/acpe","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1704.04081","atlas_url":"https://app.syntology.ai/?focus=1704.04081","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}