{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/triply-supervised-decoder-networks-for-joint","title":"Triply Supervised Decoder Networks for Joint Detection and Segmentation","arxiv_id":"1809.09299","date":"2018-09-25","proceeding":"CVPR 2019 6","authors":["Jiale Cao","Yanwei Pang","Xuelong. Li"],"abstract":"Joint object detection and semantic segmentation can be applied to many\nfields, such as self-driving cars and unmanned surface vessels. An initial and\nimportant progress towards this goal has been achieved by simply sharing the\ndeep convolutional features for the two tasks. However, this simple scheme is\nunable to make full use of the fact that detection and segmentation are\nmutually beneficial. To overcome this drawback, we propose a framework called\nTripleNet where triple supervisions including detection-oriented supervision,\nclass-aware segmentation supervision, and class-agnostic segmentation\nsupervision are imposed on each layer of the decoder network. Class-agnostic\nsegmentation supervision provides an objectness prior knowledge for both\nsemantic segmentation and object detection. Besides the three types of\nsupervisions, two light-weight modules (i.e., inner-connected module and\nattention skip-layer fusion) are also incorporated into each layer of the\ndecoder. In the proposed framework, detection and segmentation can sufficiently\nboost each other. Moreover, class-agnostic and class-aware segmentation on each\ndecoder layer are not performed at the test stage. Therefore, no extra\ncomputational costs are introduced at the test stage. Experimental results on\nthe VOC2007 and VOC2012 datasets demonstrate that the proposed TripleNet is\nable to improve both the detection and segmentation accuracies without adding\nextra computational costs.","url_abs":"http://arxiv.org/abs/1809.09299v1","url_pdf":"http://arxiv.org/pdf/1809.09299v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"self-driving-cars","task_name":"Self-Driving Cars"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/semantic-segmentation-on-pascal-voc-2012","task":"Semantic Segmentation","dataset":"PASCAL VOC 2012 test","model":"TripleNet","rank_in_archive_order":18,"of":51,"metrics":{"Mean IoU":"83.3%"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}