Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable
Rounak Saha, Gurusha Juneja, Dayita Chaudhuri, Naveeja Sajeevan, Nihar Shah, Danish Pruthi
Abstract
A number of scientific conferences and journals have recently enacted policies that prohibit LLM usage by peer reviewers, except for polishing, paraphrasing, and grammar correction of otherwise human-written reviews. But, are these policies enforceable? To answer this question, we assemble a dataset of peer reviews simulating multiple levels of human-AI collaboration, and evaluate five state-of-the-art detectors, including two commercial systems. Our analysis shows that all detectors misclassify a non-trivial fraction of LLM-polished reviews as AI-generated, thereby risking false accusations of academic misconduct. We further investigate whether peer-review-specific signals, including access to the paper manuscript and the constrained domain of scientific writing, can be leveraged to improve detection. While incorporating such signals yields measurable gains in some settings, we identify limitations in each approach and find that none meets the accuracy standards required for identifying AI use in peer reviews. Importantly, our results suggest that recent public estimates of AI use in peer reviews through the use of AI-text detectors should be interpreted with caution, as current detectors misclassify mixed reviews (collaborative human-AI outputs) as fully AI generated, potentially overstating the extent of policy violations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9e61112d-816a-43b4-b518-10c5a4ffffceCited by top-tier papers1
Ask how each one uses itBuilds on12
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning et al.ICML 2023 · 988 citations
- RADAR: Robust AI-Text Detection via Adversarial LearningXiaomeng Hu, Pin-Yu Chen, Tsung-Yi HoNeurIPS 2023 · 315 citations
- Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability CurvatureGuangsheng Bao, Yanbin Zhao, Zhiyang Teng, Linyi Yang et al.ICLR 2024 · 311 citations
- Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated TextAbhimanyu Hans, Avi Schwarzschild, Valeriia Cherepanova, Hamid Kazemi et al.ICML 2024 · 262 citations
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer ReviewsWeixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp et al.ICML 2024 · 213 citations
Related papers
- Sem-Detect: Semantic Level Detection of AI Generated Peer-ReviewsAndré Duarte, Brian Tufts, Aditya Oke, Fei Fang et al.ICML 2026 · 1 citation
- Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer ReviewSungduk Yu, Man Luo, Avinash Madasu, Vasudev Lal et al.ICLR 2026 · 24 citations
- CoCoNUTS: Concentrating on Content while Neglecting Uninformative Textual Styles for AI-Generated Peer Review DetectionYihan Chen, Jiawei Chen, Guozhao Mo, Xuanang Chen et al.ACL 2026 · 1 citation
- AI Wrote My Paper and All I Got was This False Negative:* Measuring the Efficacy of Commercial AI Text DetectorsSeth Layton, Bernardo B. P. Medeiros, Kevin R. B. Butler, Patrick TraynorS&P 2026 · 3 citations
- The AI Review Lottery: Widespread AI-Assisted Peer Reviews Boost Paper Scores and Acceptance RatesGiuseppe Russo, Manoel Horta Ribeiro, Tim R. Davidson, Veniamin Veselovsky et al.CSCW 2025 · 10 citations
