All Accept, No Reject: Evaluating LLMs as "Peer" Reviewers
Nitin Verma, Asheley R. Landrum
Abstract
An exponential rise in manuscript submission volume has strained the peer review system, prompting interest in automation from overburdened scholars and publishers. We systematically evaluated GPT-4.1, GPT-4o, o1, o3, o3-mini, and GPT-5 as “peer” reviewers, comparing their evaluation and acceptance of 137 manuscripts from an open dataset (PeerRead) with the corresponding human-generated reviews. While o3’s and GPT-5’s acceptance rates were close to the human benchmark (∼ 67% of submissions), others approved nearly every paper (>98%); all models performed extremely poorly on accuracy, precision, and recall metrics. To probe this striking “yes-bias”, we profiled the LLMs using Schwartz’s Portrait Values Questionnaire (PVQ-RR) and found that all LLMs emphasized self-transcendence and openness-to-change and de-emphasized conservation and self-enhancement. We argue that value orientations of LLMs we investigated are misaligned with the values underpinning peer review, and suggest new research on aligning AI judgment systems with human goals in this context.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- The AI Review Lottery: Widespread AI-Assisted Peer Reviews Boost Paper Scores and Acceptance RatesGiuseppe Russo, Manoel Horta Ribeiro, Tim R. Davidson, Veniamin Veselovsky et al.CSCW 2025 · 10 citations
- LLM or Human? Perceptions of Trust and Quality in Research SummariesNil-Jana Akpinar, Sandeep Avula, Chia-Jung Lee, Brandon Dang et al.CHI 2026 · 2 citations
- Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not EnforceableRounak Saha, Gurusha Juneja, Dayita Chaudhuri, Naveeja Sajeevan et al.ICML 2026 · 3 citations
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer ReviewsWeixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp et al.ICML 2024 · 213 citations
- Sem-Detect: Semantic Level Detection of AI Generated Peer-ReviewsAndré Duarte, Brian Tufts, Aditya Oke, Fei Fang et al.ICML 2026 · 1 citation
