{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/holistic-instance-level-human-parsing","title":"Holistic, Instance-Level Human Parsing","arxiv_id":"1709.03612","date":"2017-09-11","proceeding":null,"authors":["Qizhu Li","Anurag Arnab","Philip H. S. Torr"],"abstract":"Object parsing -- the task of decomposing an object into its semantic parts\n-- has traditionally been formulated as a category-level segmentation problem.\nConsequently, when there are multiple objects in an image, current methods\ncannot count the number of objects in the scene, nor can they determine which\npart belongs to which object. We address this problem by segmenting the parts\nof objects at an instance-level, such that each pixel in the image is assigned\na part label, as well as the identity of the object it belongs to. Moreover, we\nshow how this approach benefits us in obtaining segmentations at coarser\ngranularities as well. Our proposed network is trained end-to-end given\ndetections, and begins with a category-level segmentation module. Thereafter, a\ndifferentiable Conditional Random Field, defined over a variable number of\ninstances for every input image, reasons about the identity of each part by\nassociating it with a human detection. In contrast to other approaches, our\nmethod can handle the varying number of people in each image and our holistic\nnetwork produces state-of-the-art results in instance-level part and human\nsegmentation, together with competitive results in category-level part\nsegmentation, all achieved by a single forward-pass through our neural network.","url_abs":"http://arxiv.org/abs/1709.03612v1","url_pdf":"http://arxiv.org/pdf/1709.03612v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"holistic-instance-level-human-parsing","repo_url":"https://github.com/torrvision/caffe-tvg","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"human-detection","task_name":"Human Detection"},{"task_slug":"human-parsing","task_name":"Human Parsing"},{"task_slug":"multi-human-parsing","task_name":"Multi-Human Parsing"},{"task_slug":"object","task_name":"Object"},{"task_slug":"segmentation","task_name":"Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/multi-human-parsing-on-pascal-person-part","task":"Multi-Human Parsing","dataset":"PASCAL-Part","model":"Holistic instance-level","rank_in_archive_order":2,"of":3,"metrics":{"AP 0.5":"40.60%"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1709.03612","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}