Benchmarking Egocentric Clinical Intent Understanding Capability for Medical Multimodal Large Language Models
Shaonan Liu, Guo Yu, Xiaoling Luo, Shiyi Zheng, Jie Liu, Wenting Chen, Linlin Shen
Abstract
Medical Multimodal Large Language Models (Med-MLLMs) require egocentric clinical intent understanding for real-world deployment, yet existing benchmarks fail to evaluate this critical capability. To address these challenges, we introduce MedGaze-Bench, the first benchmark leveraging clinician gaze as a Cognitive Cursor to assess intent understanding across surgery, emergency simulation, and diagnostic interpretation. Our benchmark addresses three fundamental challenges: visual homogeneity of anatomical structures, strict temporal-causal dependencies in clinical workflows, and implicit adherence to safety protocols. We propose a Three-Dimensional Clinical Intent Framework evaluating: (1) Spatial Intent: discriminating precise targets amid visual noise, (2) Temporal Intent: inferring causal rationale through retrospective and prospective reasoning, and (3) Standard Intent: verifying protocol compliance through safety checks. Beyond accuracy metrics, we introduce Trap QA mechanisms to stress-test clinical reliability by penalizing hallucinations and cognitive sycophancy. Experiments reveal current MLLMs struggle with egocentric intent due to over-reliance on global features, leading to fabricated observations and uncritical acceptance of invalid instructions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- A Large-scale Study of Spatiotemporal Representation Learning with a New Benchmark on Action RecognitionAndong Deng, Taojiannan Yang, Chen ChenICCV 2023 · 18 citations
- EgoLife: Towards Egocentric Life AssistantJingkang Yang, Shuai Liu, Hongming Guo, Yuhao Dong et al.CVPR 2025
- ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition BenchmarkRonghao Dang, Yuqian Yuan, Wenqi Zhang, Yifei Xin et al.CVPR 2025
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video AnalysisChaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li et al.CVPR 2025
- EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question AnsweringSheng Zhou, Junbin Xiao, Qingyun Li, Yicong Li et al.CVPR 2025
Related papers
- Beyond Single View: A Comprehensive Benchmark for Medical Multimodal Large Language Models on Multi-Image UnderstandingDexuan Xu, Jiayin Yuan, Jianing Wang, Yanyuan Chen et al.ACL 2026
- Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language ModelsWenxuan Wang, Zizhan Ma, Guo Yu, Yiu-Fai Cheung et al.ACL 2026 · 9 citations
- SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought BenchmarkGui Wang, YongSong Zhou, Kaijun Deng, Wooi Ping Cheah et al.CVPR 2026
- MultiMedBench: A Scenario-Aware Benchmark for Evaluating Knowledge Editing in Medical VQAShengtao Wen, Haodong Chen, Yadong Wang, Zhongying Pan et al.AAAI 2026
- M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image UnderstandingJuntao Jiang, Jiangning Zhang, Yali Bi, Jinsheng Bai et al.ICLR 2026 · 3 citations
