Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
Weixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp, Hancheng Cao, Xuandong Zhao, Lingjiao Chen, Haotian Ye, Sheng Liu, Zhi Huang, Daniel A. McFarland, James Y. Zou
Abstract
We present an approach for estimating the fraction of text in a large corpus which is likely to be substantially modified or produced by a large language model (LLM). Our maximum likelihood model leverages expert-written and AI-generated reference texts to accurately and efficiently examine real-world LLM-use at the corpus level. We apply this approach to a case study of scientific peer review in AI conferences that took place after the release of ChatGPT: ICLR 2024, NeurIPS 2023, CoRL 2023 and EMNLP 2023. Our results suggest that between 6.5% and 16.9% of text submitted as peer reviews to these conferences could have been substantially modified by LLMs, i.e. beyond spell-checking or minor writing updates. The circumstances in which generated text occurs offer insight into user behavior: the estimated fraction of LLM-generated text is higher in reviews which report lower confidence, were submitted close to the deadline, and from reviewers who are less likely to respond to author rebuttals. We also observe corpus-level trends in generated text which may be too subtle to detect at the individual level, and discuss the implications of such trends on peer review. We call for future interdisciplinary work to examine how LLM use is changing our information and knowledge practices.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8dfe2341-8473-47e8-bf7b-b294e08c1199Cited by top-tier papers20
- Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language ModelsJiayu Wang, Yifei Ming, Zhenmei Shi, Vibhav Vineet et al.NeurIPS 2024 · 166 citations
- DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking ProcessMinjun Zhu, Yixuan Weng, Linyi Yang, Yue ZhangACL 2025 · 70 citations
- To Rely or Not to Rely? Evaluating Interventions for Appropriate Reliance on Large Language ModelsJessica Y. Bo, Sophia Wan, Ashton AndersonCHI 2025 · 31 citations
- Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer ReviewSungduk Yu, Man Luo, Avinash Madasu, Vasudev Lal et al.ICLR 2026 · 24 citations
- AgentReview: Exploring Peer Review Dynamics with LLM AgentsYiqiao Jin, Qinlin Zhao, Yiyang Wang, Hao Chen et al.EMNLP 2024 · 24 citations
Builds on17
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning et al.ICML 2023 · 988 citations
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
- RADAR: Robust AI-Text Detection via Adversarial LearningXiaomeng Hu, Pin-Yu Chen, Tsung-Yi HoNeurIPS 2023 · 315 citations
- Provable Robust Watermarking for AI-Generated TextXuandong Zhao, Prabhanjan Vijendra Ananth, Lei Li, Yu-Xiang WangICLR 2024 · 312 citations
- The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context LearningBill Yuchen Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri et al.ICLR 2024 · 299 citations
Related papers
- The AI Review Lottery: Widespread AI-Assisted Peer Reviews Boost Paper Scores and Acceptance RatesGiuseppe Russo, Manoel Horta Ribeiro, Tim R. Davidson, Veniamin Veselovsky et al.CSCW 2025 · 10 citations
- Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not EnforceableRounak Saha, Gurusha Juneja, Dayita Chaudhuri, Naveeja Sajeevan et al.ICML 2026 · 3 citations
- Sem-Detect: Semantic Level Detection of AI Generated Peer-ReviewsAndré Duarte, Brian Tufts, Aditya Oke, Fei Fang et al.ICML 2026 · 1 citation
- 'Quis custodiet ipsos custodes?' Who will watch the watchmen? On Detecting AI-generated peer-reviewsSandeep Kumar, Mohit Sahu, Vardhan Gacche, Tirthankar Ghosal et al.EMNLP 2024 · 4 citations
- People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated textJenna Russell, Marzena Karpinska, Mohit IyyerACL 2025 · 39 citations
