{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sipmask-spatial-information-preservation-for","title":"SipMask: Spatial Information Preservation for Fast Image and Video Instance Segmentation","arxiv_id":"2007.14772","date":"2020-07-29","proceeding":"ECCV 2020 8","authors":["Jiale Cao","Rao Muhammad Anwer","Hisham Cholakkal","Fahad Shahbaz Khan","Yanwei Pang","Ling Shao"],"abstract":"Single-stage instance segmentation approaches have recently gained popularity due to their speed and simplicity, but are still lagging behind in accuracy, compared to two-stage methods. We propose a fast single-stage instance segmentation method, called SipMask, that preserves instance-specific spatial information by separating mask prediction of an instance to different sub-regions of a detected bounding-box. Our main contribution is a novel light-weight spatial preservation (SP) module that generates a separate set of spatial coefficients for each sub-region within a bounding-box, leading to improved mask predictions. It also enables accurate delineation of spatially adjacent instances. Further, we introduce a mask alignment weighting loss and a feature alignment scheme to better correlate mask prediction with object detection. On COCO test-dev, our SipMask outperforms the existing single-stage methods. Compared to the state-of-the-art single-stage TensorMask, SipMask obtains an absolute gain of 1.0% (mask AP), while providing a four-fold speedup. In terms of real-time capabilities, SipMask outperforms YOLACT with an absolute gain of 3.0% (mask AP) under similar settings, while operating at comparable speed on a Titan Xp. We also evaluate our SipMask for real-time video instance segmentation, achieving promising results on YouTube-VIS dataset. The source code is available at https://github.com/JialeCao001/SipMask.","url_abs":"https://arxiv.org/abs/2007.14772v1","url_pdf":"https://arxiv.org/pdf/2007.14772v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"sipmask-spatial-information-preservation-for","repo_url":"https://github.com/JialeCao001/SipMask","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"instance-segmentation","task_name":"Instance Segmentation"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"real-time-instance-segmentation","task_name":"Real-time Instance Segmentation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"video-instance-segmentation","task_name":"Video Instance Segmentation"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/instance-segmentation-on-coco","task":"Instance Segmentation","dataset":"COCO test-dev","model":"SipMask (ResNet-101, single-scale test)","rank_in_archive_order":87,"of":112,"metrics":{"AP50":"60.2","AP75":"40.8","APL":"54.3","APM":"40.8","APS":"17.8","mask AP":"38.1"},"uses_additional_data":false},{"leaderboard":"/sota/real-time-instance-segmentation-on-mscoco","task":"Real-time Instance Segmentation","dataset":"MSCOCO","model":"SipMask++ (ResNet-101, single-scale test)","rank_in_archive_order":11,"of":22,"metrics":{"AP50":"55.6","AP75":"37.6","APL":"56.8","APM":"38.3","APS":"11.2","Frame (fps)":"27.0 (Titan Xp)","mask AP":"35.4"},"uses_additional_data":false},{"leaderboard":"/sota/real-time-instance-segmentation-on-mscoco","task":"Real-time Instance Segmentation","dataset":"MSCOCO","model":"SipMask (ResNet-101, single-scale test)","rank_in_archive_order":17,"of":22,"metrics":{"AP50":"53.4","AP75":"34.3","APL":"54.0","APM":"35.6","APS":"9.3","Frame (fps)":"31.3 (Titan Xp)","mask AP":"32.8"},"uses_additional_data":false},{"leaderboard":"/sota/real-time-instance-segmentation-on-mscoco","task":"Real-time Instance Segmentation","dataset":"MSCOCO","model":"SipMask (ResNet-50, single-scale test)","rank_in_archive_order":18,"of":22,"metrics":{"AP50":"51.9","AP75":"32.3","APL":"49.8","APM":"33.6","APS":"9.2","Frame (fps)":"41.7 (Titan Xp)","mask AP":"31.2"},"uses_additional_data":false},{"leaderboard":"/sota/video-instance-segmentation-on-youtube-vis-1","task":"Video Instance Segmentation","dataset":"YouTube-VIS validation","model":"SipMask (ResNet-50, ms-train, single-scale test)","rank_in_archive_order":36,"of":44,"metrics":{"AP50":"54.1","AP75":"35.8","AR1":"35.4","AR10":"40.1","mask AP":"33.7"},"uses_additional_data":false},{"leaderboard":"/sota/video-instance-segmentation-on-youtube-vis-1","task":"Video Instance Segmentation","dataset":"YouTube-VIS validation","model":"SipMask (ResNet-50, single-scale test)","rank_in_archive_order":38,"of":44,"metrics":{"AP50":"53","AP75":"33.3","AR1":"33.5","AR10":"38.9","mask AP":"32.5"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2007.14772","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}