{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/knowledge-guided-deep-fractal-neural-networks","title":"Knowledge-Guided Deep Fractal Neural Networks for Human Pose Estimation","arxiv_id":"1705.02407","date":"2017-05-05","proceeding":null,"authors":["Guanghan Ning","Zhi Zhang","Zhihai He"],"abstract":"Human pose estimation using deep neural networks aims to map input images\nwith large variations into multiple body keypoints which must satisfy a set of\ngeometric constraints and inter-dependency imposed by the human body model.\nThis is a very challenging nonlinear manifold learning process in a very high\ndimensional feature space. We believe that the deep neural network, which is\ninherently an algebraic computation system, is not the most effecient way to\ncapture highly sophisticated human knowledge, for example those highly coupled\ngeometric characteristics and interdependence between keypoints in human poses.\nIn this work, we propose to explore how external knowledge can be effectively\nrepresented and injected into the deep neural networks to guide its training\nprocess using learned projections that impose proper prior. Specifically, we\nuse the stacked hourglass design and inception-resnet module to construct a\nfractal network to regress human pose images into heatmaps with no explicit\ngraphical modeling. We encode external knowledge with visual features which are\nable to characterize the constraints of human body models and evaluate the\nfitness of intermediate network output. We then inject these external features\ninto the neural network using a projection matrix learned using an auxiliary\ncost function. The effectiveness of the proposed inception-resnet module and\nthe benefit in guided learning with knowledge projection is evaluated on two\nwidely used benchmarks. Our approach achieves state-of-the-art performance on\nboth datasets.","url_abs":"http://arxiv.org/abs/1705.02407v2","url_pdf":"http://arxiv.org/pdf/1705.02407v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"knowledge-guided-deep-fractal-neural-networks","repo_url":"https://github.com/Guanghan/GNet-pose","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"pose-estimation","task_name":"Pose Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/pose-estimation-on-leeds-sports-poses","task":"Pose Estimation","dataset":"Leeds Sports Poses","model":"Stacked hourglass + Inception-resnet","rank_in_archive_order":7,"of":18,"metrics":{"PCK":"93.9%"},"uses_additional_data":false},{"leaderboard":"/sota/pose-estimation-on-mpii-human-pose","task":"Pose Estimation","dataset":"MPII Human Pose","model":"Stacked hourglass + Inception-resnet","rank_in_archive_order":20,"of":46,"metrics":{"PCKh-0.5":"91.2"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1705.02407","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}