{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/masklab-instance-segmentation-by-refining","title":"MaskLab: Instance Segmentation by Refining Object Detection with Semantic and Direction Features","arxiv_id":"1712.04837","date":"2017-12-13","proceeding":"CVPR 2018 6","authors":["Liang-Chieh Chen","Alexander Hermans","George Papandreou","Florian Schroff","Peng Wang","Hartwig Adam"],"abstract":"In this work, we tackle the problem of instance segmentation, the task of\nsimultaneously solving object detection and semantic segmentation. Towards this\ngoal, we present a model, called MaskLab, which produces three outputs: box\ndetection, semantic segmentation, and direction prediction. Building on top of\nthe Faster-RCNN object detector, the predicted boxes provide accurate\nlocalization of object instances. Within each region of interest, MaskLab\nperforms foreground/background segmentation by combining semantic and direction\nprediction. Semantic segmentation assists the model in distinguishing between\nobjects of different semantic classes including background, while the direction\nprediction, estimating each pixel's direction towards its corresponding center,\nallows separating instances of the same semantic class. Moreover, we explore\nthe effect of incorporating recent successful methods from both segmentation\nand detection (i.e. atrous convolution and hypercolumn). Our proposed model is\nevaluated on the COCO instance segmentation benchmark and shows comparable\nperformance with other state-of-art models.","url_abs":"http://arxiv.org/abs/1712.04837v1","url_pdf":"http://arxiv.org/pdf/1712.04837v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"instance-segmentation","task_name":"Instance Segmentation"},{"task_slug":"object","task_name":"Object"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"prediction","task_name":"Prediction"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[{"method_slug":"1x1-convolution","method_name":"1x1 Convolution"},{"method_slug":"average-pooling","method_name":"Average Pooling"},{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"bottleneck-residual-block","method_name":"Bottleneck Residual Block"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"dilated-convolution","method_name":"Dilated Convolution"},{"method_slug":"global-average-pooling","method_name":"Global Average Pooling"},{"method_slug":"kaiming-initialization","method_name":"Kaiming Initialization"},{"method_slug":"max-pooling","method_name":"Max Pooling"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"residual-block","method_name":"Residual Block"},{"method_slug":"residual-connection","method_name":"Residual Connection"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/instance-segmentation-on-coco","task":"Instance Segmentation","dataset":"COCO test-dev","model":"MaskLab+ (ResNet-101, JFT)","rank_in_archive_order":88,"of":112,"metrics":{"mask AP":"38.1%"},"uses_additional_data":true}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1712.04837","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}