{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/detect-and-track-efficient-pose-estimation-in","title":"Detect-and-Track: Efficient Pose Estimation in Videos","arxiv_id":"1712.09184","date":"2017-12-26","proceeding":"CVPR 2018 6","authors":["Rohit Girdhar","Georgia Gkioxari","Lorenzo Torresani","Manohar Paluri","Du Tran"],"abstract":"This paper addresses the problem of estimating and tracking human body\nkeypoints in complex, multi-person video. We propose an extremely lightweight\nyet highly effective approach that builds upon the latest advancements in human\ndetection and video understanding. Our method operates in two-stages: keypoint\nestimation in frames or short clips, followed by lightweight tracking to\ngenerate keypoint predictions linked over the entire video. For frame-level\npose estimation we experiment with Mask R-CNN, as well as our own proposed 3D\nextension of this model, which leverages temporal information over small clips\nto generate more robust frame predictions. We conduct extensive ablative\nexperiments on the newly released multi-person video pose estimation benchmark,\nPoseTrack, to validate various design choices of our model. Our approach\nachieves an accuracy of 55.2% on the validation and 51.8% on the test set using\nthe Multi-Object Tracking Accuracy (MOTA) metric, and achieves state of the art\nperformance on the ICCV 2017 PoseTrack keypoint tracking challenge.","url_abs":"http://arxiv.org/abs/1712.09184v2","url_pdf":"http://arxiv.org/pdf/1712.09184v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"detect-and-track-efficient-pose-estimation-in","repo_url":"https://github.com/facebookresearch/DetectAndTrack","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"caffe2","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"human-detection","task_name":"Human Detection"},{"task_slug":"keypoint-estimation","task_name":"Keypoint Estimation"},{"task_slug":"multi-object-tracking","task_name":"Multi-Object Tracking"},{"task_slug":"object-tracking","task_name":"Object Tracking"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"pose-tracking","task_name":"Pose Tracking"},{"task_slug":"video-understanding","task_name":"Video Understanding"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"mask-r-cnn","method_name":"Mask R-CNN"},{"method_slug":"rpn","method_name":"RPN"},{"method_slug":"roi-align","method_name":"RoIAlign"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/keypoint-detection-on-coco-test-challenge","task":"Keypoint Detection","dataset":"COCO test-challenge","model":"Girdhar et al.","rank_in_archive_order":7,"of":8,"metrics":{"AR":"70.2","ARM":"60.7"},"uses_additional_data":false},{"leaderboard":"/sota/pose-tracking-on-posetrack2017","task":"Pose Tracking","dataset":"PoseTrack2017","model":"ProTracker","rank_in_archive_order":8,"of":10,"metrics":{"MOTA":"51.82","mAP":"59.56"},"uses_additional_data":true}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1712.09184","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}