{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/simultaneous-detection-and-segmentation","title":"Simultaneous Detection and Segmentation","arxiv_id":"1407.1808","date":"2014-07-07","proceeding":null,"authors":["Bharath Hariharan","Pablo Arbeláez","Ross Girshick","Jitendra Malik"],"abstract":"We aim to detect all instances of a category in an image and, for each\ninstance, mark the pixels that belong to it. We call this task Simultaneous\nDetection and Segmentation (SDS). Unlike classical bounding box detection, SDS\nrequires a segmentation and not just a box. Unlike classical semantic\nsegmentation, we require individual object instances. We build on recent work\nthat uses convolutional neural networks to classify category-independent region\nproposals (R-CNN [16]), introducing a novel architecture tailored for SDS. We\nthen use category-specific, top- down figure-ground predictions to refine our\nbottom-up proposals. We show a 7 point boost (16% relative) over our baselines\non SDS, a 5 point boost (10% relative) over state-of-the-art on semantic\nsegmentation, and state-of-the-art performance in object detection. Finally, we\nprovide diagnostic tools that unpack performance and provide directions for\nfuture work.","url_abs":"http://arxiv.org/abs/1407.1808v1","url_pdf":"http://arxiv.org/pdf/1407.1808v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"diagnostic","task_name":"Diagnostic"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/object-detection-on-pascal-voc-2012","task":"Object Detection","dataset":"PASCAL VOC 2012","model":"SDS","rank_in_archive_order":5,"of":7,"metrics":{"MAP":"50.7"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-pascal-voc-2012","task":"Semantic Segmentation","dataset":"PASCAL VOC 2012 test","model":"CK","rank_in_archive_order":51,"of":51,"metrics":{"Mean IoU":"51.6%"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1407.1808","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}