{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deeperlab-single-shot-image-parser","title":"DeeperLab: Single-Shot Image Parser","arxiv_id":"1902.05093","date":"2019-02-13","proceeding":null,"authors":["Tien-Ju Yang","Maxwell D. Collins","Yukun Zhu","Jyh-Jing Hwang","Ting Liu","Xiao Zhang","Vivienne Sze","George Papandreou","Liang-Chieh Chen"],"abstract":"We present a single-shot, bottom-up approach for whole image parsing. Whole\nimage parsing, also known as Panoptic Segmentation, generalizes the tasks of\nsemantic segmentation for 'stuff' classes and instance segmentation for 'thing'\nclasses, assigning both semantic and instance labels to every pixel in an\nimage. Recent approaches to whole image parsing typically employ separate\nstandalone modules for the constituent semantic and instance segmentation tasks\nand require multiple passes of inference. Instead, the proposed DeeperLab image\nparser performs whole image parsing with a significantly simpler, fully\nconvolutional approach that jointly addresses the semantic and instance\nsegmentation tasks in a single-shot manner, resulting in a streamlined system\nthat better lends itself to fast processing. For quantitative evaluation, we\nuse both the instance-based Panoptic Quality (PQ) metric and the proposed\nregion-based Parsing Covering (PC) metric, which better captures the image\nparsing quality on 'stuff' classes and larger object instances. We report\nexperimental results on the challenging Mapillary Vistas dataset, in which our\nsingle model achieves 31.95% (val) / 31.6% PQ (test) and 55.26% PC (val) with 3\nframes per second (fps) on GPU or near real-time speed (22.6 fps on GPU) with\nreduced accuracy.","url_abs":"http://arxiv.org/abs/1902.05093v2","url_pdf":"http://arxiv.org/pdf/1902.05093v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":null,"task_name":"GPU"},{"task_slug":"instance-segmentation","task_name":"Instance Segmentation"},{"task_slug":"panoptic-segmentation","task_name":"Panoptic Segmentation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[{"method_slug":"speed","method_name":"SPEED"},{"method_slug":"pc","method_name":"pc"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/panoptic-segmentation-on-cityscapes-val","task":"Panoptic Segmentation","dataset":"Cityscapes val","model":"DeeperLab (Xception-71)","rank_in_archive_order":33,"of":37,"metrics":{"PQ":"56.5"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1902.05093","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}