{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/understanding-humans-in-crowded-scenes-deep","title":"Understanding Humans in Crowded Scenes: Deep Nested Adversarial Learning and A New Benchmark for Multi-Human Parsing","arxiv_id":"1804.03287","date":"2018-04-10","proceeding":null,"authors":["Jian Zhao","Jianshu Li","Yu Cheng","Li Zhou","Terence Sim","Shuicheng Yan","Jiashi Feng"],"abstract":"Despite the noticeable progress in perceptual tasks like detection, instance\nsegmentation and human parsing, computers still perform unsatisfactorily on\nvisually understanding humans in crowded scenes, such as group behavior\nanalysis, person re-identification and autonomous driving, etc. To this end,\nmodels need to comprehensively perceive the semantic information and the\ndifferences between instances in a multi-human image, which is recently defined\nas the multi-human parsing task. In this paper, we present a new large-scale\ndatabase \"Multi-Human Parsing (MHP)\" for algorithm development and evaluation,\nand advances the state-of-the-art in understanding humans in crowded scenes.\nMHP contains 25,403 elaborately annotated images with 58 fine-grained semantic\ncategory labels, involving 2-26 persons per image and captured in real-world\nscenes from various viewpoints, poses, occlusion, interactions and background.\nWe further propose a novel deep Nested Adversarial Network (NAN) model for\nmulti-human parsing. NAN consists of three Generative Adversarial Network\n(GAN)-like sub-nets, respectively performing semantic saliency prediction,\ninstance-agnostic parsing and instance-aware clustering. These sub-nets form a\nnested structure and are carefully designed to learn jointly in an end-to-end\nway. NAN consistently outperforms existing state-of-the-art solutions on our\nMHP and several other datasets, and serves as a strong baseline to drive the\nfuture research for multi-human parsing.","url_abs":"http://arxiv.org/abs/1804.03287v3","url_pdf":"http://arxiv.org/pdf/1804.03287v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"understanding-humans-in-crowded-scenes-deep","repo_url":"https://github.com/ZhaoJ9014/Multi-Human-Parsing","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"understanding-humans-in-crowded-scenes-deep","repo_url":"https://github.com/open-mmlab/mmpose","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"autonomous-driving","task_name":"Autonomous Driving"},{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":null,"task_name":"Generative Adversarial Network"},{"task_slug":"human-parsing","task_name":"Human Parsing"},{"task_slug":"instance-segmentation","task_name":"Instance Segmentation"},{"task_slug":"multi-human-parsing","task_name":"Multi-Human Parsing"},{"task_slug":"person-re-identification","task_name":"Person Re-Identification"},{"task_slug":"saliency-prediction","task_name":"Saliency Prediction"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/multi-human-parsing-on-mhp-v10","task":"Multi-Human Parsing","dataset":"MHP v1.0","model":"NAN","rank_in_archive_order":1,"of":4,"metrics":{"AP 0.5":"57.09%"},"uses_additional_data":false},{"leaderboard":"/sota/multi-human-parsing-on-mhp-v20","task":"Multi-Human Parsing","dataset":"MHP v2.0","model":"NAN","rank_in_archive_order":3,"of":5,"metrics":{"AP 0.5":"25.14%"},"uses_additional_data":false},{"leaderboard":"/sota/multi-human-parsing-on-pascal-person-part","task":"Multi-Human Parsing","dataset":"PASCAL-Part","model":"NAN","rank_in_archive_order":1,"of":3,"metrics":{"AP 0.5":"59.70%"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1804.03287","atlas_url":"https://app.syntology.ai/?focus=1804.03287","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}