{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-human-parsing-with-active-template","title":"Deep Human Parsing with Active Template Regression","arxiv_id":"1503.02391","date":"2015-03-09","proceeding":null,"authors":["Xiaodan Liang","Si Liu","Xiaohui Shen","Jianchao Yang","Luoqi Liu","Jian Dong","Liang Lin","Shuicheng Yan"],"abstract":"In this work, the human parsing task, namely decomposing a human image into\nsemantic fashion/body regions, is formulated as an Active Template Regression\n(ATR) problem, where the normalized mask of each fashion/body item is expressed\nas the linear combination of the learned mask templates, and then morphed to a\nmore precise mask with the active shape parameters, including position, scale\nand visibility of each semantic region. The mask template coefficients and the\nactive shape parameters together can generate the human parsing results, and\nare thus called the structure outputs for human parsing. The deep Convolutional\nNeural Network (CNN) is utilized to build the end-to-end relation between the\ninput human image and the structure outputs for human parsing. More\nspecifically, the structure outputs are predicted by two separate networks. The\nfirst CNN network is with max-pooling, and designed to predict the template\ncoefficients for each label mask, while the second CNN network is without\nmax-pooling to preserve sensitivity to label mask position and accurately\npredict the active shape parameters. For a new image, the structure outputs of\nthe two networks are fused to generate the probability of each label for each\npixel, and super-pixel smoothing is finally used to refine the human parsing\nresult. Comprehensive evaluations on a large dataset well demonstrate the\nsignificant superiority of the ATR framework over other state-of-the-arts for\nhuman parsing. In particular, the F1-score reaches $64.38\\%$ by our ATR\nframework, significantly higher than $44.76\\%$ based on the state-of-the-art\nalgorithm.","url_abs":"http://arxiv.org/abs/1503.02391v1","url_pdf":"http://arxiv.org/pdf/1503.02391v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deep-human-parsing-with-active-template","repo_url":"https://github.com/AemikaChow/DATASOURCE","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"human-parsing","task_name":"Human Parsing"},{"task_slug":null,"task_name":"Position"},{"task_slug":"regression-1","task_name":"regression"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1503.02391","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}