{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/bi-directional-cross-modality-feature","title":"Bi-directional Cross-Modality Feature Propagation with Separation-and-Aggregation Gate for RGB-D Semantic Segmentation","arxiv_id":"2007.09183","date":"2020-07-17","proceeding":"ECCV 2020 8","authors":["Xiaokang Chen","Kwan-Yee Lin","Jingbo Wang","Wayne Wu","Chen Qian","Hongsheng Li","Gang Zeng"],"abstract":"Depth information has proven to be a useful cue in the semantic segmentation of RGB-D images for providing a geometric counterpart to the RGB representation. Most existing works simply assume that depth measurements are accurate and well-aligned with the RGB pixels and models the problem as a cross-modal feature fusion to obtain better feature representations to achieve more accurate segmentation. This, however, may not lead to satisfactory results as actual depth data are generally noisy, which might worsen the accuracy as the networks go deeper. In this paper, we propose a unified and efficient Cross-modality Guided Encoder to not only effectively recalibrate RGB feature responses, but also to distill accurate depth information via multiple stages and aggregate the two recalibrated representations alternatively. The key of the proposed architecture is a novel Separation-and-Aggregation Gating operation that jointly filters and recalibrates both representations before cross-modality aggregation. Meanwhile, a Bi-direction Multi-step Propagation strategy is introduced, on the one hand, to help to propagate and fuse information between the two modalities, and on the other hand, to preserve their specificity along the long-term propagation process. Besides, our proposed encoder can be easily injected into the previous encoder-decoder structures to boost their performance on RGB-D semantic segmentation. Our model outperforms state-of-the-arts consistently on both in-door and out-door challenging datasets. Code of this work is available at https://charlescxk.github.io/","url_abs":"https://arxiv.org/abs/2007.09183v1","url_pdf":"https://arxiv.org/pdf/2007.09183v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"bi-directional-cross-modality-feature","repo_url":"https://github.com/David-zaiwang/114_rgbd_seg","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"bi-directional-cross-modality-feature","repo_url":"https://github.com/charlesCXK/RGBD_Semantic_Segmentation_PyTorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"specificity","task_name":"Specificity"},{"task_slug":"thermal-image-segmentation","task_name":"Thermal Image Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/object-detection-on-dsec","task":"Object Detection","dataset":"DSEC","model":"SAGate","rank_in_archive_order":11,"of":12,"metrics":{"mAP":"19.6"},"uses_additional_data":false},{"leaderboard":"/sota/object-detection-on-pku-ddd17-car","task":"Object Detection","dataset":"PKU-DDD17-Car","model":"SAGate","rank_in_archive_order":6,"of":14,"metrics":{"mAP50":"82.0"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-bjroad","task":"Semantic Segmentation","dataset":"BJRoad","model":"SA-Gate","rank_in_archive_order":4,"of":11,"metrics":{"IoU":"62.14"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-event-based","task":"Semantic Segmentation","dataset":"Event-based Segmentation Dataset","model":"SA-Gate","rank_in_archive_order":3,"of":6,"metrics":{"mIoU":"84.08"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-eventscape","task":"Semantic Segmentation","dataset":"EventScape","model":"SA-Gate","rank_in_archive_order":5,"of":12,"metrics":{"mIoU":"53.94"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-llrgbd-synthetic","task":"Semantic Segmentation","dataset":"LLRGBD-synthetic","model":"SA-Gate (ResNet-101)","rank_in_archive_order":8,"of":8,"metrics":{"mIoU":"61.79"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-nyu-depth-v2","task":"Semantic Segmentation","dataset":"NYU Depth v2","model":"SA-Gate","rank_in_archive_order":45,"of":121,"metrics":{"Mean IoU":"52.4%"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-porto","task":"Semantic Segmentation","dataset":"Porto","model":"SA-Gate","rank_in_archive_order":4,"of":6,"metrics":{"IoU":"72.21"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-potsdam","task":"Semantic Segmentation","dataset":"Potsdam","model":"SA-Gate","rank_in_archive_order":8,"of":11,"metrics":{"mIoU":"84.28"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-sun-rgbd","task":"Semantic Segmentation","dataset":"SUN-RGBD","model":"TokenFusion (Ti)","rank_in_archive_order":24,"of":44,"metrics":{"Mean IoU":"49.4%"},"uses_additional_data":true},{"leaderboard":"/sota/semantic-segmentation-on-thud-robotic-dataset","task":"Semantic Segmentation","dataset":"THUD Robotic Dataset","model":"SA-Gate","rank_in_archive_order":1,"of":4,"metrics":{"mIoU":"83.19"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-tlcgis","task":"Semantic Segmentation","dataset":"TLCGIS","model":"SA-Gate","rank_in_archive_order":1,"of":6,"metrics":{"IoU":"84.20"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-us3d","task":"Semantic Segmentation","dataset":"US3D","model":"SA-Gate","rank_in_archive_order":4,"of":11,"metrics":{"mIoU":"83.62"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-urbanlf","task":"Semantic Segmentation","dataset":"UrbanLF","model":"SA-Gate","rank_in_archive_order":4,"of":14,"metrics":{"mIoU (Real)":"n.a.","mIoU (Syn)":"79.53"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-vaihingen","task":"Semantic Segmentation","dataset":"Vaihingen","model":"SA-Gate","rank_in_archive_order":3,"of":13,"metrics":{"mIoU":"81.03"},"uses_additional_data":false},{"leaderboard":"/sota/thermal-image-segmentation-on-mfn-dataset","task":"Thermal Image Segmentation","dataset":"MFN Dataset","model":"SA-Gate","rank_in_archive_order":48,"of":55,"metrics":{"mIOU":"45.8"},"uses_additional_data":false},{"leaderboard":"/sota/thermal-image-segmentation-on-noisy-rs-rgb-t","task":"Thermal Image Segmentation","dataset":"Noisy RS RGB-T Dataset","model":"SA-Gate","rank_in_archive_order":4,"of":6,"metrics":{"mIoU":"54.0"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2007.09183","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}