{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/look-into-person-joint-body-parsing-pose","title":"Look into Person: Joint Body Parsing & Pose Estimation Network and A New Benchmark","arxiv_id":"1804.01984","date":"2018-04-05","proceeding":null,"authors":["Xiaodan Liang","Ke Gong","Xiaohui Shen","Liang Lin"],"abstract":"Human parsing and pose estimation have recently received considerable\ninterest due to their substantial application potentials. However, the existing\ndatasets have limited numbers of images and annotations and lack a variety of\nhuman appearances and coverage of challenging cases in unconstrained\nenvironments. In this paper, we introduce a new benchmark named \"Look into\nPerson (LIP)\" that provides a significant advancement in terms of scalability,\ndiversity, and difficulty, which are crucial for future developments in\nhuman-centric analysis. This comprehensive dataset contains over 50,000\nelaborately annotated images with 19 semantic part labels and 16 body joints,\nwhich are captured from a broad range of viewpoints, occlusions, and background\ncomplexities. Using these rich annotations, we perform detailed analyses of the\nleading human parsing and pose estimation approaches, thereby obtaining\ninsights into the successes and failures of these methods. To further explore\nand take advantage of the semantic correlation of these two tasks, we propose a\nnovel joint human parsing and pose estimation network to explore efficient\ncontext modeling, which can simultaneously predict parsing and pose with\nextremely high quality. Furthermore, we simplify the network to solve human\nparsing by exploring a novel self-supervised structure-sensitive learning\napproach, which imposes human pose structures into the parsing results without\nresorting to extra supervision. The dataset, code and models are available at\nhttp://www.sysu-hcp.net/lip/.","url_abs":"http://arxiv.org/abs/1804.01984v1","url_pdf":"http://arxiv.org/pdf/1804.01984v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"look-into-person-joint-body-parsing-pose","repo_url":"https://github.com/Engineering-Course/LIP_JPPNet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"look-into-person-joint-body-parsing-pose","repo_url":"https://github.com/Engineering-Course/LIP_SSL","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"look-into-person-joint-body-parsing-pose","repo_url":"https://github.com/andrewjong/SwapNet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"human-parsing","task_name":"Human Parsing"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[{"method_slug":"1x1-convolution","method_name":"1x1 Convolution"},{"method_slug":"average-pooling","method_name":"Average Pooling"},{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"bottleneck-residual-block","method_name":"Bottleneck Residual Block"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"global-average-pooling","method_name":"Global Average Pooling"},{"method_slug":"kaiming-initialization","method_name":"Kaiming Initialization"},{"method_slug":"max-pooling","method_name":"Max Pooling"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"residual-block","method_name":"Residual Block"},{"method_slug":"residual-connection","method_name":"Residual Connection"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/semantic-segmentation-on-lip-val","task":"Semantic Segmentation","dataset":"LIP val","model":"JPPNet (ResNet-101)","rank_in_archive_order":10,"of":13,"metrics":{"mIoU":"51.37%"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1804.01984","atlas_url":"https://app.syntology.ai/?focus=1804.01984","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}