{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/flowing-convnets-for-human-pose-estimation-in","title":"Flowing ConvNets for Human Pose Estimation in Videos","arxiv_id":"1506.02897","date":"2015-06-09","proceeding":"ICCV 2015 12","authors":["Tomas Pfister","James Charles","Andrew Zisserman"],"abstract":"The objective of this work is human pose estimation in videos, where multiple\nframes are available. We investigate a ConvNet architecture that is able to\nbenefit from temporal context by combining information across the multiple\nframes using optical flow.\n  To this end we propose a network architecture with the following novelties:\n(i) a deeper network than previously investigated for regressing heatmaps; (ii)\nspatial fusion layers that learn an implicit spatial model; (iii) optical flow\nis used to align heatmap predictions from neighbouring frames; and (iv) a final\nparametric pooling layer which learns to combine the aligned heatmaps into a\npooled confidence map.\n  We show that this architecture outperforms a number of others, including one\nthat uses optical flow solely at the input layers, one that regresses joint\ncoordinates directly, and one that predicts heatmaps without spatial fusion.\n  The new architecture outperforms the state of the art by a large margin on\nthree video pose estimation datasets, including the very challenging Poses in\nthe Wild dataset, and outperforms other deep methods that don't use a graphical\nmodel on the single-image FLIC benchmark (and also Chen & Yuille and Tompson et\nal. in the high precision region).","url_abs":"http://arxiv.org/abs/1506.02897v2","url_pdf":"http://arxiv.org/pdf/1506.02897v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"flowing-convnets-for-human-pose-estimation-in","repo_url":"https://github.com/parthapratimbanik/facial-keypoint-detection-udacity-ppb","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"}],"methods":[{"method_slug":"heatmap","method_name":"Heatmap"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1506.02897","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}