SciMDR: Advancing Scientific Multimodal Document Reasoning
Ziyu Chen, Yilun Zhao, Chengye Wang, Rilyn Han, Manasi Patwardhan, Arman Cohan
Abstract
Constructing scientific multimodal document reasoning datasets for foundation model training involves an inherent trade-off among scale, faithfulness, and realism. To address this challenge, we introduce the synthesize-and-reground framework, a two-stage pipeline comprising: (1) Claim-Centric QA Synthesis, which generates faithful, isolated QA pairs and reasoning on focused segments, and (2) Document-Scale Regrounding, which programmatically re-embeds these pairs into full-document tasks to ensure realistic complexity. Using this framework, we construct SCIMDR, a large-scale training dataset for cross-modal comprehension, comprising 300K QA pairs with explicit reasoning chains across 20K scientific papers. We further construct SCIMDR-EVAL, an expert-annotated benchmark to evaluate multimodal comprehension within full-length scientific workflows. Experiments demonstrate that models fine-tuned on SCIMDR achieve significant improvements across multiple scientific QA benchmarks (e.g., ChartQA, SPIQA, SCIMDR-EVAL), particularly in those tasks requiring complex document-level reasoning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6f206562-e759-4846-a582-9bcc351f79afBuilds on13
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang et al.ICML 2024 · 1,191 citations
- LongBench: A Bilingual, Multitask Benchmark for Long Context UnderstandingYushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu et al.ACL 2024 · 94 citations
- QASA: Advanced Question Answering on Scientific ArticlesYoonjoo Lee, Kyungjae Lee, Sunghyun Park, Dasol Hwang et al.ICML 2023 · 76 citations
- Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research PapersZhijian Xu, Yilun Zhao, Manasi Patwardhan, Lovekesh Vig et al.ACL 2025 · 17 citations
Related papers
- SciDQA: A Deep Reading Comprehension Dataset over Scientific PapersShruti Singh, Nandan Sarkar, Arman CohanEMNLP 2024 · 2 citations
- DomainCQA: Crafting Knowledge-Intensive QA from Domain-Specific ChartsYujing Lu, Ling Zhong, Jing Yang, Weiming Li et al.AAAI 2026
- Long Grounded Thoughts: Synthesizing Visual Problems and Reasoning Chains at ScaleDavid Acuna, Chao-Han Huck Yang, Yuntian Deng, Jaehun Jung et al.ICML 2026 · 1 citation
- CRIT: Graph-Based Automatic Data Synthesis to Enhance Cross-Modal Multi-Hop ReasoningJunyoung Sung, Seungwoo Lyu, Minjun Kim, Sumin An et al.CVPR 2026 · 2 citations
- SciVer: Evaluating Foundation Models for Multimodal Scientific Claim VerificationChengye Wang, Yifei Shen, Zexi Kuang, Arman Cohan et al.ACL 2025 · 8 citations
