Navigating Through Paper Flood: Advancing LLM-Based Paper Evaluation Through Domain-Aware Retrieval and Latent Reasoning
Wuqiang Zheng, Yiyan Xu, Xinyu Lin, Chongming Gao, Wenjie Wang, Fuli Feng
Abstract
With the rapid and continuous increase in academic publications, identifying high-quality research has become an increasingly pressing challenge. While recent methods leveraging Large Language Models (LLMs) for automated paper evaluation have shown great promise, they are often constrained by outdated domain knowledge and limited reasoning capabilities. In this work, we present PaperEval, a novel LLM-based framework for automated paper evaluation that addresses these limitations through two key components: 1) a domain-aware paper retrieval module that retrieves relevant concurrent work to support contextualized assessments of novelty and contributions, and 2) a latent reasoning mechanism that enables deep understanding of complex motivations and methodologies, along with comprehensive comparison against concurrently related work, to support more accurate and reliable evaluation. To guide the reasoning process, we introduce a progressive ranking optimization strategy that encourages the LLM to iteratively refine its predictions with an emphasis on relative comparison. Experiments on two datasets demonstrate that PaperEval consistently outperforms existing methods in both academic impact and paper quality evaluation. In addition, we deploy PaperEval in a real-world paper recommendation system for filtering high-quality papers, which has gained strong engagement on social media---amassing over 8,000 subscribers and attracting over 10,000 views for many filtered high-quality papers---demonstrating the practical effectiveness of PaperEval.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 04b79bec-3be3-4208-9a03-04e0ac78ac1dBuilds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Efficient Reasoning with Hidden ThinkingXuan Shen, Yizhou Wang, Yufa Zhou, Xiangxi Shi et al.ICML 2026 · 56 citations
- HINTS: Citation Time Series Prediction for New Publications via Dynamic Heterogeneous Information Network EmbeddingSong Jiang, Bernard Koch, Yizhou SunWWW 2021 · 45 citations
- Neural Reasoning Networks: Efficient Interpretable Neural Networks with Automatic Textual ExplanationsStephen Carrow, Kyle Erwin, Olga Vilenskaia, Parikshit Ram et al.AAAI 2025 · 4 citations
Related papers
- DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking ProcessMinjun Zhu, Yixuan Weng, Linyi Yang, Yue ZhangACL 2025 · 70 citations
- From Replication to Redesign: Exploring Pairwise Comparisons for LLM-Based Peer ReviewYaohui Zhang, Haijing Zhang, Wenlong Ji, Tianyu Hua et al.NeurIPS 2025 · 15 citations
- P2P: Automated Paper-to-Poster Generation and Fine-Grained BenchmarkTao Sun, Enhao Pan, Zhengkai Yang, Kaixin Sui et al.ICLR 2026 · 19 citations
- Auto-PRE: An Automatic and Cost-Efficient Peer-Review Framework for Language Generation EvaluationJunjie Chen, Weihang Su, Zhumin Chu, Haitao Li et al.AAAI 2026
- Mind the Blind Spots: A Focus-Level Evaluation Framework for LLM ReviewsHyungyu Shin, Jingyu Tang, Yoonjoo Lee, Nayoung Kim et al.EMNLP 2025 · 2 citations
