Papers › Missing Modality Robustness in Semi-Supervised Multi-Modal Semantic Segmentation
Missing Modality Robustness in Semi-Supervised Multi-Modal Semantic Segmentation
Harsh Maheshwari, Yen-Cheng Liu, Zsolt Kira
Using multiple spatial modalities has been proven helpful in improving semantic segmentation performance. However, there are several real-world challenges that have yet to be addressed: (a) improving label efficiency and (b) enhancing robustness in realistic scenarios where modalities are missing at the test time. To address these challenges, we first propose a simple yet efficient multi-modal fusion mechanism Linear Fusion, that performs better than the state-of-the-art multi-modal models even with limited supervision. Second, we propose M3L: Multi-modal Teacher for Masked Modality Learning, a semi-supervised framework that not only improves the multi-modal performance but also makes the model robust to the realistic missing modality scenario using unlabeled data. We create the first benchmark for semi-supervised multi-modal semantic segmentation and also report the robustness to missing modalities. Our proposal shows an absolute improvement of up to 10% on robust mIoU above the most competitive baselines. Our code is available at https://github.com/harshm121/M3L
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
1 archive task tag without a task page not shown.
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Semantic Segmentation | SUN-RGBD | DFormer-L | Mean IoU (test) | 48.17 | #44 of 44 | Archive leaderboard | report |
| Semantic Segmentation | Stanford2D3D - RGBD | Linear Fusion (Segformer B2) | mIoU | 57.16 | #4 of 6 | Archive leaderboard | report |
| Semi-Supervised Semantic Segmentation | 2D-3D-S | M3L (Linear Fusion B2) | mIoU (0.1% labels) | 40.05 | #1 of 1 | Archive leaderboard | report |
| Semi-Supervised Semantic Segmentation | 2D-3D-S | M3L (Linear Fusion B2) | mIoU (0.2% labels) | 44.62 | #1 of 1 | Archive leaderboard | report |
| Semi-Supervised Semantic Segmentation | 2D-3D-S | M3L (Linear Fusion B2) | mIoU (1% labels) | 49.28 | #1 of 1 | Archive leaderboard | report |
| Semi-Supervised Semantic Segmentation | Stanford 2D-3D | M3L (Linear Fusion - Segformer B2) | MM-Robust mIoU (0.1% labels) | 41.36 | #1 of 2 | Archive leaderboard | report |
| Semi-Supervised Semantic Segmentation | Stanford 2D-3D | M3L (Linear Fusion - Segformer B2) | mIoU (0.1% labels) | 44.1 | #1 of 2 | Archive leaderboard | report |
| Semi-Supervised Semantic Segmentation | Stanford 2D-3D | Mean Teacher (Linear Fusion - Segformer B2) | mIoU (0.1% labels) | 41.7 | #2 of 2 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Introduced by this paper: M3L
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections