{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/missing-modality-robustness-in-semi","title":"Missing Modality Robustness in Semi-Supervised Multi-Modal Semantic Segmentation","arxiv_id":"2304.10756","date":"2023-04-21","proceeding":null,"authors":["Harsh Maheshwari","Yen-Cheng Liu","Zsolt Kira"],"abstract":"Using multiple spatial modalities has been proven helpful in improving semantic segmentation performance. However, there are several real-world challenges that have yet to be addressed: (a) improving label efficiency and (b) enhancing robustness in realistic scenarios where modalities are missing at the test time. To address these challenges, we first propose a simple yet efficient multi-modal fusion mechanism Linear Fusion, that performs better than the state-of-the-art multi-modal models even with limited supervision. Second, we propose M3L: Multi-modal Teacher for Masked Modality Learning, a semi-supervised framework that not only improves the multi-modal performance but also makes the model robust to the realistic missing modality scenario using unlabeled data. We create the first benchmark for semi-supervised multi-modal semantic segmentation and also report the robustness to missing modalities. Our proposal shows an absolute improvement of up to 10% on robust mIoU above the most competitive baselines. Our code is available at https://github.com/harshm121/M3L","url_abs":"https://arxiv.org/abs/2304.10756v1","url_pdf":"https://arxiv.org/pdf/2304.10756v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"missing-modality-robustness-in-semi","repo_url":"https://github.com/harshm121/m3l","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":null,"task_name":"RGBD Semantic Segmentation"},{"task_slug":"robust-semi-supervised-rgbd-semantic","task_name":"Robust Semi-Supervised RGBD Semantic Segmentation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"semi-supervised-rgbd-semantic-segmentation","task_name":"Semi-Supervised RGBD Semantic Segmentation"},{"task_slug":"semi-supervised-semantic-segmentation","task_name":"Semi-Supervised Semantic Segmentation"}],"methods":[{"method_slug":"m3l","method_name":"M3L"},{"method_slug":"test","method_name":"Test"}],"datasets_introduced":[],"methods_introduced":[{"slug":"m3l","name":"M3L","full_name":"Multi-modal Teacher for Masked Modality Learning"}],"results":[{"leaderboard":"/sota/semantic-segmentation-on-sun-rgbd","task":"Semantic Segmentation","dataset":"SUN-RGBD","model":"DFormer-L","rank_in_archive_order":44,"of":44,"metrics":{"Mean IoU (test)":"48.17"},"uses_additional_data":true},{"leaderboard":"/sota/semantic-segmentation-on-stanford2d3d-rgbd","task":"Semantic Segmentation","dataset":"Stanford2D3D - RGBD","model":"Linear Fusion (Segformer B2)","rank_in_archive_order":4,"of":6,"metrics":{"mIoU":"57.16"},"uses_additional_data":false},{"leaderboard":"/sota/semi-supervised-semantic-segmentation-on-2d","task":"Semi-Supervised Semantic Segmentation","dataset":"2D-3D-S","model":"M3L (Linear Fusion B2)","rank_in_archive_order":1,"of":1,"metrics":{"mIoU (0.1% labels)":"40.05","mIoU (0.2% labels)":"44.62","mIoU (1% labels)":"49.28"},"uses_additional_data":false},{"leaderboard":"/sota/semi-supervised-semantic-segmentation-on-32","task":"Semi-Supervised Semantic Segmentation","dataset":"Stanford 2D-3D","model":"M3L (Linear Fusion - Segformer B2)","rank_in_archive_order":1,"of":2,"metrics":{"MM-Robust mIoU (0.1% labels)":"41.36","mIoU (0.1% labels)":"44.1"},"uses_additional_data":false},{"leaderboard":"/sota/semi-supervised-semantic-segmentation-on-32","task":"Semi-Supervised Semantic Segmentation","dataset":"Stanford 2D-3D","model":"Mean Teacher (Linear Fusion - Segformer B2)","rank_in_archive_order":2,"of":2,"metrics":{"mIoU (0.1% labels)":"41.7"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2304.10756","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}