DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process
Minjun Zhu, Yixuan Weng, Linyi Yang, Yue Zhang
摘要
Large Language Models (LLMs) are increasingly utilized in scientific research assessment, particularly in automated paper review. However, existing LLM-based review systems face significant challenges, including limited domain expertise, hallucinated reasoning, and a lack of structured evaluation. To address these limitations, we introduce DeepReview, a multi-stage framework designed to emulate expert reviewers by incorporating structured analysis, literature retrieval, and evidence-based argumentation. Using DeepReview-13K, a curated dataset with structured annotations, we train DeepReviewer-14B, which outperforms CycleReviewer-70B with fewer tokens. In its best mode, DeepReviewer-14B achieves win rates of 88.21% and 80.20% against GPT-o1 and DeepSeek-R1 in evaluations. Our work sets a new benchmark for LLM-based paper review, with all resources publicly available. The code, model, dataset and demo have be released in http://ai-researcher.net.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research IdeasChenglei Si, Tatsunori Hashimoto, Diyi YangICLR 2026 · 被引用 60 次
- DeepScientist: Advancing Frontier-Pushing Scientific Findings ProgressivelyYixuan Weng, Minjun Zhu, Qiujie Xie, Qiyao Sun 等ICLR 2026 · 被引用 57 次
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi 等EMNLP 2025 · 被引用 37 次
- Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer ReviewSungduk Yu, Man Luo, Avinash Madasu, Vasudev Lal 等ICLR 2026 · 被引用 24 次
- From Replication to Redesign: Exploring Pairwise Comparisons for LLM-Based Peer ReviewYaohui Zhang, Haijing Zhang, Wenlong Ji, Tianyu Hua 等NeurIPS 2025 · 被引用 15 次
它引用的顶会 Paper19
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
相关 Paper
- DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report GenerationJanghoon Han, Heegyu Kim, Changho Lee, Dahm Lee 等ICML 2026 · 被引用 9 次
- CycleResearcher: Improving Automated Research via Automated ReviewYixuan Weng, Minjun Zhu, Guangsheng Bao, Hongbo Zhang 等ICLR 2025
- Navigating Through Paper Flood: Advancing LLM-Based Paper Evaluation Through Domain-Aware Retrieval and Latent ReasoningWuqiang Zheng, Yiyan Xu, Xinyu Lin, Chongming Gao 等AAAI 2026
- Agent Reviewers: Domain-specific Multimodal Agents with Shared Memory for Paper ReviewKai Lu, Shixiong Xu, Jinqiu Li, Kun Ding 等ICML 2025
- Can Large Language Models Match the Conclusions of Systematic Reviews?Christopher Polzak, Alejandro Lozano, Min Woo Sun, James Burgess 等ICLR 2026 · 被引用 9 次
