{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multi-context-attention-for-human-pose","title":"Multi-Context Attention for Human Pose Estimation","arxiv_id":"1702.07432","date":"2017-02-24","proceeding":"CVPR 2017 7","authors":["Xiao Chu","Wei Yang","Wanli Ouyang","Cheng Ma","Alan L. Yuille","Xiaogang Wang"],"abstract":"In this paper, we propose to incorporate convolutional neural networks with a\nmulti-context attention mechanism into an end-to-end framework for human pose\nestimation. We adopt stacked hourglass networks to generate attention maps from\nfeatures at multiple resolutions with various semantics. The Conditional Random\nField (CRF) is utilized to model the correlations among neighboring regions in\nthe attention map. We further combine the holistic attention model, which\nfocuses on the global consistency of the full human body, and the body part\nattention model, which focuses on the detailed description for different body\nparts. Hence our model has the ability to focus on different granularity from\nlocal salient regions to global semantic-consistent spaces. Additionally, we\ndesign novel Hourglass Residual Units (HRUs) to increase the receptive field of\nthe network. These units are extensions of residual units with a side branch\nincorporating filters with larger receptive fields, hence features with various\nscales are learned and combined within the HRUs. The effectiveness of the\nproposed multi-context attention mechanism and the hourglass residual units is\nevaluated on two widely used human pose estimation benchmarks. Our approach\noutperforms all existing methods on both benchmarks over all the body parts.","url_abs":"http://arxiv.org/abs/1702.07432v1","url_pdf":"http://arxiv.org/pdf/1702.07432v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multi-context-attention-for-human-pose","repo_url":"https://github.com/bearpaw/pose-attention","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"torch","reach":null},{"paper_slug":"multi-context-attention-for-human-pose","repo_url":"https://github.com/wbenbihi/hourglasstensorlfow","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"pose-estimation","task_name":"Pose Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/pose-estimation-on-leeds-sports-poses","task":"Pose Estimation","dataset":"Leeds Sports Poses","model":"Multi-Context Attention","rank_in_archive_order":8,"of":18,"metrics":{"PCK":"92.6%"},"uses_additional_data":false},{"leaderboard":"/sota/pose-estimation-on-mpii-human-pose","task":"Pose Estimation","dataset":"MPII Human Pose","model":"Multi-Context Attention","rank_in_archive_order":17,"of":46,"metrics":{"PCKh-0.5":"91.5"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1702.07432","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}