SciDQA: A Deep Reading Comprehension Dataset over Scientific Papers
Shruti Singh, Nandan Sarkar, Arman Cohan
摘要
Scientific literature is typically dense, requiring significant background knowledge and deep comprehension for effective engagement. We introduce SciDQA, a new dataset for reading comprehension that challenges language models to deeply understand scientific articles, consisting of 2,937 QA pairs. Unlike other scientific QA datasets, SciDQA sources questions from peer reviews by domain experts and answers by paper authors, ensuring a thorough examination of the literature. We enhance the dataset’s quality through a process that carefully decontextualizes the content, tracks the source document across different versions, and incorporates a bibliography for multi-document question-answering. Questions in SciDQA necessitate reasoning across figures, tables, equations, appendices, and supplementary materials, and require multi-document reasoning. We evaluate several open-source and proprietary LLMs across various configurations to explore their capabilities in generating relevant and factual responses, as opposed to simple review memorization. Our comprehensive evaluation, based on metrics for surface-level and semantic similarity, highlights notable performance discrepancies. SciDQA represents a rigorously curated, naturally derived scientific QA dataset, designed to facilitate research on complex reasoning within the domain of question answering for scientific texts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal InconsistenciesLukas Selch, Yufang Hou, Muhammad Jehanzeb Mirza, Sivan Doveh 等ICLR 2026 · 被引用 2 次
- AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level EvaluationTiancheng Huang, Ruisheng Cao, Yuxin Zhang, Zhangyi Kang 等ICLR 2026 · 被引用 1 次
- SciMDR: Advancing Scientific Multimodal Document ReasoningZiyu Chen, Yilun Zhao, Chengye Wang, Rilyn Han 等ACL 2026 · 被引用 1 次
- LECTOR: Joint Learning of Scientific Reasoning Graphs and Introduction GenerationJiabei Xiao, Yizhou Wang, Chen Tang, Pengze Li 等ICML 2026
- NeuSym-RAG: Hybrid Neural Symbolic Retrieval with Multiview Structuring for PDF Question AnsweringRuisheng Cao, Hanchong Zhang, Tiancheng Huang, Zhangyi Kang 等ACL 2025
它引用的顶会 Paper6
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Nougat: Neural Optical Understanding for Academic DocumentsLukas Blecher, Guillem Cucurull, Thomas Scialom, Robert StojnicICLR 2024 · 被引用 243 次
- QASA: Advanced Question Answering on Scientific ArticlesYoonjoo Lee, Kyungjae Lee, Sunghyun Park, Dasol Hwang 等ICML 2023 · 被引用 76 次
- Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human EvaluationYixin Liu, Alexander R. Fabbri, Pengfei Liu, Yilun Zhao 等ACL 2023 · 被引用 50 次
- Relatedly: Scaffolding Literature Reviews with Existing Related Work SectionsSrishti Palani, Aakanksha Naik, Doug Downey, Amy X. Zhang 等CHI 2023 · 被引用 37 次
相关 Paper
- A Multi-Task Learning Framework for Reading Comprehension of Scientific Tabular DataXu Yang, Meihui Zhang, Ju Fan, Zeyu Luo 等ICDE 2024 · 被引用 1 次
- SciRIFF: A Resource to Enhance Language Model Instruction-Following over Scientific LiteratureDavid Wadden, Kejian Shi, Jacob Morrison, Alan Li 等EMNLP 2025 · 被引用 2 次
- M³-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question AnsweringJiatong Ma, Longteng Guo, Yuchen Liu, Zijia Zhao 等ACL 2026
- MMQA: Evaluating LLMs with Multi-Table Multi-Hop Complex QuestionsJian Wu, Linyi Yang, Dongyuan Li, Yuliang Ji 等ICLR 2025
- VisualMRC: Machine Reading Comprehension on Document ImagesRyota Tanaka, Kyosuke Nishida, Sen YoshidaAAAI 2021 · 被引用 201 次
