The AI Review Lottery: Widespread AI-Assisted Peer Reviews Boost Paper Scores and Acceptance Rates
Giuseppe Russo, Manoel Horta Ribeiro, Tim R. Davidson, Veniamin Veselovsky, Robert West
Abstract
Journals and conferences worry that peer reviews assisted by artificial intelligence (AI), in particular, large language models (LLMs), may negatively influence the validity and fairness of the peer-review system, a cornerstone of modern science. In this work, we address this concern with a study of the prevalence and impact of AI-assisted peer reviews in the context of the 2024 International Conference on Learning Representations (ICLR), a large and prestigious machine-learning conference. Our contributions are threefold. Firstly, we obtain a lower bound for the prevalence of AI-assisted reviews at ICLR 2024 using the closed- and open-source LLM detectors, estimating that at least 15.8% of reviews were written with AI assistance. Secondly, we estimate the impact of AI-assisted reviews on submission scores. Considering pairs of reviews with different scores assigned to the same paper, we find that in 53.4% of pairs, the AI-assisted review scores higher than the human review (p = 0.002; relative difference in probability of scoring higher: +14.4% in favor of AI-assisted reviews). Thirdly, we assess the impact of receiving an AI-assisted peer review on submission acceptance. In a matched study, submissions near the acceptance threshold that received an AI-assisted peer review were 4.9 percentage points (p = 0.024) more likely to be accepted than submissions that did not. Overall, we show that AI-assisted reviews are consequential to the peer-review process and offer a discussion on future implications of current trends.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking ProcessMinjun Zhu, Yixuan Weng, Linyi Yang, Yue ZhangACL 2025 · 70 citations
- Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer ReviewSungduk Yu, Man Luo, Avinash Madasu, Vasudev Lal et al.ICLR 2026 · 24 citations
- Valid Survey Simulations with Limited Human Data: The Roles of Prompting, Fine-Tuning, and RectificationStefan Krsteski, Giuseppe Russo, Serina Chang, Robert West et al.ACL 2026 · 10 citations
- Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not EnforceableRounak Saha, Gurusha Juneja, Dayita Chaudhuri, Naveeja Sajeevan et al.ICML 2026 · 3 citations
- Mind the Blind Spots: A Focus-Level Evaluation Framework for LLM ReviewsHyungyu Shin, Jingyu Tang, Yoonjoo Lee, Nayoung Kim et al.EMNLP 2025 · 2 citations
Builds on3
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 865 citations
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer ReviewsWeixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp et al.ICML 2024 · 213 citations
Related papers
- Sem-Detect: Semantic Level Detection of AI Generated Peer-ReviewsAndré Duarte, Brian Tufts, Aditya Oke, Fei Fang et al.ICML 2026 · 1 citation
- All Accept, No Reject: Evaluating LLMs as "Peer" ReviewersNitin Verma, Asheley R. LandrumCHI 2026 · 2 citations
- LLM or Human? Perceptions of Trust and Quality in Research SummariesNil-Jana Akpinar, Sandeep Avula, Chia-Jung Lee, Brandon Dang et al.CHI 2026 · 2 citations
- CoCoNUTS: Concentrating on Content while Neglecting Uninformative Textual Styles for AI-Generated Peer Review DetectionYihan Chen, Jiawei Chen, Guozhao Mo, Xuanang Chen et al.ACL 2026 · 1 citation
- NAIPv2: Debiased Pairwise Learning for Efficient Paper Quality EstimationPenghai Zhao, Jinyu Tian, Qinghua Xing, Xin Zhang et al.ICLR 2026 · 6 citations
