Papers › MM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of...

MM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical Environments

4 Mar 2025CVPR 2025 1arXiv:2503.02579archive 2025-07-28

Ege Özsoy, Chantal Pellegrini, Tobias Czempiel, Felix Tristram, Kun Yuan, David Bani-Harouni, Ulrich Eck, Benjamin Busam, Matthias Keicher, Nassir Navab

Operating rooms (ORs) are complex, high-stakes environments requiring precise understanding of interactions among medical staff, tools, and equipment for enhancing surgical assistance, situational awareness, and patient safety. Current datasets fall short in scale, realism and do not capture the multimodal nature of OR scenes, limiting progress in OR modeling. To this end, we introduce MM-OR, a realistic and large-scale multimodal spatiotemporal OR dataset, and the first dataset to enable multimodal scene graph generation. MM-OR captures comprehensive OR scenes containing RGB-D data, detail views, audio, speech transcripts, robotic logs, and tracking data and is annotated with panoptic segmentations, semantic scene graphs, and downstream task labels. Further, we propose MM2SG, the first multimodal large vision-language model for scene graph generation, and through extensive experiments, demonstrate its ability to effectively leverage multimodal inputs. Together, MM-OR and MM2SG establish a new benchmark for holistic OR understanding, and open the path towards multimodal scene analysis in complex, high-stakes environments. Our code, and data is available at https://github.com/egeozsoy/MM-OR.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

egeozsoy/MM-OR officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

2D Panoptic SegmentationGraph GenerationLanguage ModelingLanguage ModellingScene Graph GenerationVideo Panoptic Segmentation

Datasets

Introduced by this paper, per the archive.

MM-OR

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
2D Panoptic Segmentation 4D-OR MM-OR VPQ 71.8 #1 of 1 Archive leaderboard report
2D Panoptic Segmentation MM-OR MM-OR VPQ 67.5 #1 of 1 Archive leaderboard report
Scene Graph Generation 4D-OR MM2SG F1 0.901 #2 of 5 Archive leaderboard report
Scene Graph Generation MM-OR MM2SG Macro F1 0.529 #1 of 1 Archive leaderboard report
Video Panoptic Segmentation 4D-OR MM-OR-VPQ4 VPQ 69.8 #1 of 2 Archive leaderboard report
Video Panoptic Segmentation 4D-OR MM-OR-VPQ8 VPQ 69.2 #2 of 2 Archive leaderboard report
Video Panoptic Segmentation MM-OR MM-OR-VPQ4 VPQ 67.0 #1 of 2 Archive leaderboard report
Video Panoptic Segmentation MM-OR MM-OR-VPQ8 VPQ 66.4 #2 of 2 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections