MM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical Environments
Ege Özsoy, Chantal Pellegrini, Tobias Czempiel, Felix Tristram, Kun Yuan, David Bani-Harouni, Ulrich Eck, Benjamin Busam, Matthias Keicher, Nassir Navab
Abstract
Multiview RGB-D Video (5 cameras) Detail RGB Videos (3 cameras) Low Exposure RGB Video Point Cloud Sound Robot Screen, Tracker Data and Logs Panoptic Segmentations Scene Graphs Downstream Tasks Robot Phase: Robot Preparation Complete head surgeon operating table sa wi ng saw holdi ng patient robot ma nip ula tin g mps lying on nurse instrument table preparing c lo s e to Next Action: Hammering Sterility Breach: No Figure 1. Overview of a single timepoint in MM-OR, illustrating the multimodal data provided for each sample: RGB-D video from multiple angles, detailed RGB views, low-exposure video, point cloud data, robot screen and tracker logs, audio and speech transcripts, panoptic segmentations, semantic scene graphs, and downstream task annotations such as robot phase, next action, and sterility breach status.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fc83ebce-53d2-491e-8a46-8a03e497a427Cited by top-tier papers4
- Towards Dynamic 3D Reconstruction of Hand-Instrument Interaction in Ophthalmic SurgeryMing Hu, Zhengdi Yu, Feilong Tang, Kaiwen Chen et al.NeurIPS 2025 · 1 citation
- DisPOSE: Projected Polystochastic Diffusion for Self-Supervised Multi-View 3D Human Pose EstimationTony Danjun Wang, Tolga Birdal, Nassir Navab, Lennart BastianICML 2026
- Modeling Long-Tail Relations in the Operating Room via In-Context Multimodal LearningBoqiang Xu, Wei Zhang, Ding Ma, Jian Liang et al.ICML 2026
- SHands: A Multi-View Dataset and Benchmark for Surgical Hand-Gesture and Error Recognition Toward Medical TrainingLe Ma, Thiago Freitas dos Santos, Nadia Magnenat-Thalmann, Katarzyna WacCVPR 2026
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
Related papers
- Bridging Vision and Language for Robust Context-Aware Surgical Point Tracking: The VL-SurgPT Dataset and BenchmarkRulin Zhou, Wenlong He, An Wang, Jianhang Zhang et al.AAAI 2026
- TeamVision: An AI-powered Learning Analytics System for Supporting Reflection in Team-based Healthcare SimulationVanessa Echeverría, Linxuan Zhao, Riordan Alfredo, Mikaela Elizabeth Milesi et al.CHI 2025 · 20 citations
- Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person PerspectivesKristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani et al.CVPR 2024
- SiM3D: Single-Instance Multiview Multimodal and Multisetup 3D Anomaly Detection BenchmarkAlex Costanzino, Pierluigi Zama Ramirez, Luigi Lella, Matteo Ragaglia et al.ICCV 2025 · 2 citations
- EgoGen: An Egocentric Synthetic Data GeneratorGen Li, Kaifeng Zhao, Siwei Zhang, Xiaozhong Lyu et al.CVPR 2024
