Benchmarking Egocentric Clinical Intent Understanding Capability for Medical Multimodal Large Language Models
Shaonan Liu, Guo Yu, Xiaoling Luo, Shiyi Zheng, Jie Liu, Wenting Chen, Linlin Shen
摘要
Medical Multimodal Large Language Models (Med-MLLMs) require egocentric clinical intent understanding for real-world deployment, yet existing benchmarks fail to evaluate this critical capability. To address these challenges, we introduce MedGaze-Bench, the first benchmark leveraging clinician gaze as a Cognitive Cursor to assess intent understanding across surgery, emergency simulation, and diagnostic interpretation. Our benchmark addresses three fundamental challenges: visual homogeneity of anatomical structures, strict temporal-causal dependencies in clinical workflows, and implicit adherence to safety protocols. We propose a Three-Dimensional Clinical Intent Framework evaluating: (1) Spatial Intent: discriminating precise targets amid visual noise, (2) Temporal Intent: inferring causal rationale through retrospective and prospective reasoning, and (3) Standard Intent: verifying protocol compliance through safety checks. Beyond accuracy metrics, we introduce Trap QA mechanisms to stress-test clinical reliability by penalizing hallucinations and cognitive sycophancy. Experiments reveal current MLLMs struggle with egocentric intent due to over-reliance on global features, leading to fabricated observations and uncritical acceptance of invalid instructions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- A Large-scale Study of Spatiotemporal Representation Learning with a New Benchmark on Action RecognitionAndong Deng, Taojiannan Yang, Chen ChenICCV 2023 · 被引用 18 次
- EgoLife: Towards Egocentric Life AssistantJingkang Yang, Shuai Liu, Hongming Guo, Yuhao Dong 等CVPR 2025
- ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition BenchmarkRonghao Dang, Yuqian Yuan, Wenqi Zhang, Yifei Xin 等CVPR 2025
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video AnalysisChaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li 等CVPR 2025
- EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question AnsweringSheng Zhou, Junbin Xiao, Qingyun Li, Yicong Li 等CVPR 2025
相关 Paper
- Beyond Single View: A Comprehensive Benchmark for Medical Multimodal Large Language Models on Multi-Image UnderstandingDexuan Xu, Jiayin Yuan, Jianing Wang, Yanyuan Chen 等ACL 2026
- Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language ModelsWenxuan Wang, Zizhan Ma, Guo Yu, Yiu-Fai Cheung 等ACL 2026 · 被引用 9 次
- SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought BenchmarkGui Wang, YongSong Zhou, Kaijun Deng, Wooi Ping Cheah 等CVPR 2026
- MultiMedBench: A Scenario-Aware Benchmark for Evaluating Knowledge Editing in Medical VQAShengtao Wen, Haodong Chen, Yadong Wang, Zhongying Pan 等AAAI 2026
- M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image UnderstandingJuntao Jiang, Jiangning Zhang, Yali Bi, Jinsheng Bai 等ICLR 2026 · 被引用 3 次
