MM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical Environments
Ege Özsoy, Chantal Pellegrini, Tobias Czempiel, Felix Tristram, Kun Yuan, David Bani-Harouni, Ulrich Eck, Benjamin Busam, Matthias Keicher, Nassir Navab
摘要
Multiview RGB-D Video (5 cameras) Detail RGB Videos (3 cameras) Low Exposure RGB Video Point Cloud Sound Robot Screen, Tracker Data and Logs Panoptic Segmentations Scene Graphs Downstream Tasks Robot Phase: Robot Preparation Complete head surgeon operating table sa wi ng saw holdi ng patient robot ma nip ula tin g mps lying on nurse instrument table preparing c lo s e to Next Action: Hammering Sterility Breach: No Figure 1. Overview of a single timepoint in MM-OR, illustrating the multimodal data provided for each sample: RGB-D video from multiple angles, detailed RGB views, low-exposure video, point cloud data, robot screen and tracker logs, audio and speech transcripts, panoptic segmentations, semantic scene graphs, and downstream task annotations such as robot phase, next action, and sterility breach status.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Towards Dynamic 3D Reconstruction of Hand-Instrument Interaction in Ophthalmic SurgeryMing Hu, Zhengdi Yu, Feilong Tang, Kaiwen Chen 等NeurIPS 2025 · 被引用 1 次
- DisPOSE: Projected Polystochastic Diffusion for Self-Supervised Multi-View 3D Human Pose EstimationTony Danjun Wang, Tolga Birdal, Nassir Navab, Lennart BastianICML 2026
- Modeling Long-Tail Relations in the Operating Room via In-Context Multimodal LearningBoqiang Xu, Wei Zhang, Ding Ma, Jian Liang 等ICML 2026
- SHands: A Multi-View Dataset and Benchmark for Surgical Hand-Gesture and Error Recognition Toward Medical TrainingLe Ma, Thiago Freitas dos Santos, Nadia Magnenat-Thalmann, Katarzyna WacCVPR 2026
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
相关 Paper
- Bridging Vision and Language for Robust Context-Aware Surgical Point Tracking: The VL-SurgPT Dataset and BenchmarkRulin Zhou, Wenlong He, An Wang, Jianhang Zhang 等AAAI 2026
- TeamVision: An AI-powered Learning Analytics System for Supporting Reflection in Team-based Healthcare SimulationVanessa Echeverría, Linxuan Zhao, Riordan Alfredo, Mikaela Elizabeth Milesi 等CHI 2025 · 被引用 20 次
- Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person PerspectivesKristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani 等CVPR 2024
- SiM3D: Single-Instance Multiview Multimodal and Multisetup 3D Anomaly Detection BenchmarkAlex Costanzino, Pierluigi Zama Ramirez, Luigi Lella, Matteo Ragaglia 等ICCV 2025 · 被引用 2 次
- EgoGen: An Egocentric Synthetic Data GeneratorGen Li, Kaifeng Zhao, Siwei Zhang, Xiaozhong Lyu 等CVPR 2024
