Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers
Zhijian Xu, Yilun Zhao, Manasi Patwardhan, Lovekesh Vig, Arman Cohan
摘要
Peer review is fundamental to scientific research, but the growing volume of publications has intensified the challenges of this expertiseintensive process. While LLMs show promise in various scientific tasks, their potential to assist with peer review, particularly in identifying paper limitations, remains understudied. We first present a comprehensive taxonomy of limitation types in scientific research, with a focus on AI. Guided by this taxonomy, for studying limitations, we present LIMITGEN, the first comprehensive benchmark for evaluating LLMs' capability to support early-stage feedback and complement human peer review. Our benchmark consists of two subsets: LIM-ITGEN-Syn, a synthetic dataset carefully created through controlled perturbations of highquality papers, and LIMITGEN-Human, a collection of real human-written limitations. To improve the ability of LLM systems to identify limitations, we augment them with literature retrieval, which is essential for grounding identifying limitations in prior scientific findings. Our approach enhances the capabilities of LLM systems to generate limitations in research papers, enabling them to provide more concrete and constructive feedback. Data yale-nlp/LimitGen Code yale-nlp/LimitGen * Equal Contributions. RQ1: How well do LLM-based systems perform in identifying limitations within scientific research? RQ2: Can RAG enhance LLMs' ability to identify limitations and provide constructive suggestions? RQ3: How can this research be applied in real-world scenarios to assist human researchers in improving their work? Read the following scientific paper and generate major limitations in this paper about its xxx.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- MemoryBench: A Benchmark for Memory and Continual Learning in LLM SystemsQingyao Ai, Yichen Tang, Changyue Wang, Jianming Long 等ICML 2026 · 被引用 47 次
- SciMDR: Advancing Scientific Multimodal Document ReasoningZiyu Chen, Yilun Zhao, Chengye Wang, Rilyn Han 等ACL 2026 · 被引用 1 次
- Automatic and Reliable Evaluation for Academic Caption-to-Figure Generation with LMMsGuanghui Ye, Huan Zhao, Qin Zhu, Fengnan Li 等ACL 2026
它引用的顶会 Paper3
- SciMON: Scientific Inspiration Machines Optimized for NoveltyQingyun Wang, Doug Downey, Heng Ji, Tom HopeACL 2024 · 被引用 22 次
- LitSearch: A Retrieval Benchmark for Scientific Literature SearchAnirudh Ajith, Mengzhou Xia, Alexis Chevalier, Tanya Goyal 等EMNLP 2024 · 被引用 8 次
- Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP ResearchersChenglei Si, Diyi Yang, Tatsunori HashimotoICLR 2025
相关 Paper
- AAAR-1.0: Assessing AI's Potential to Assist ResearchRenze Lou, Hanzi Xu, Sijia Wang, Jiangshu Du 等ICML 2025
- Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document UnderstandingSensen Gao, Shanshan Zhao, Xu Jiang, Lunhao Duan 等ACL 2026 · 被引用 7 次
- ReviewGrounder: Improving Review Substantiveness with Rubric-Guided, Tool-Integrated AgentsZhuofeng Li, Yi Lu, Dongfu Jiang, Haoxiang Zhang 等ACL 2026 · 被引用 1 次
- LLMs Assist NLP Researchers: Critique Paper (Meta-)ReviewingJiangshu Du, Yibo Wang, Wenting Zhao, Zhongfen Deng 等EMNLP 2024 · 被引用 14 次
- RExBench: Can coding agents autonomously implement AI research extensions?Nicholas Edwards, Yukyung Lee, Yujun Audrey Mao, Yulu Qin 等ACL 2026 · 被引用 10 次
