Benchmarking the Scientific Mind: A Pathology-Derived Biomedical VQA Benchmark for Complex Scientific Reasoning
Ziyu Zhao, Yiyang Liu, Yajiao Wang, Xiaotao Wang, Yang Li, Yuyang Peng, Jiaheng Zhou, Jinqiao Wang, Yingying Chen, Ge Yang, Haixin Wang
摘要
Despite progress of Multimodal Large Language Models (MLLMs) in biomedical visual question answering (VQA), existing benchmarks provide limited assessment of their scientific reasoning capabilities. Most datasets adopt single-image question construction and outcome-oriented evaluation, where correctness is judged by answer plausibility rather than alignment with experimental evidence. Such formulations fail to capture the evidence-constrained, multi-step nature of biomedical reasoning, and obscure whether models can derive conclusions through causal interpretation of experimental observations. To address these critical gaps in reasoning evaluation, we propose a principled benchmark construction framework that reconstructs scientific reasoning paths directly from biomedical literature. By jointly modeling clusters of experimentally related images together with their captions and context, the framework generates tightly coupled question–reasoning–answer triples that require multi-image integration and explicit evidence-driven inference. Based on this framework, we introduce SORBE (Scientific Observation & Reasoning for Biomedical Evaluation), a large-scale multi-image pathology-derived biomedical VQA benchmark designed to evaluate evidence alignment and multi-step experimental reasoning. Under a process-oriented evaluation metric, state-of-the-art biomedical-specialized MLLMs exhibit substantial performance degradation, revealing systematic limitations in evidence grounding and causal reasoning that are not reflected by existing benchmarks. Data and code are available at: https://github.com/UniverseOfUniverse/SORBE.git.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang 等CVPR 2024 · 被引用 213 次
- PathAsst: A Generative Foundation AI Assistant towards Artificial General Intelligence of PathologyYuxuan Sun, Chenglu Zhu, Sunyi Zheng, Kai Zhang 等AAAI 2024 · 被引用 92 次
- A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and RecommendationsMd. Tahmid Rahman Laskar, Sawsan Alqahtani, M. Saiful Bari, Mizanur Rahman 等EMNLP 2024 · 被引用 47 次
相关 Paper
- M³-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question AnsweringJiatong Ma, Longteng Guo, Yuchen Liu, Zijia Zhao 等ACL 2026
- X-PCR: A Benchmark for Cross-modality Progressive Clinical Reasoning in Ophthalmic DiagnosisGui Wang, Zehao Zhong, YongSong Zhou, Yudong Li 等CVPR 2026
- Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal ReasoningHaozhen Gong, Xiaozhong Ji, Yuansen Liu, Wenbin Wu 等CVPR 2026 · 被引用 15 次
- Beyond Single View: A Comprehensive Benchmark for Medical Multimodal Large Language Models on Multi-Image UnderstandingDexuan Xu, Jiayin Yuan, Jianing Wang, Yanyuan Chen 等ACL 2026
- M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image UnderstandingJuntao Jiang, Jiangning Zhang, Yali Bi, Jinsheng Bai 等ICLR 2026 · 被引用 3 次
