LazyReview: A Dataset for Uncovering Lazy Thinking in NLP Peer Reviews
Sukannya Purkayastha, Zhuang Li, Anne Lauscher, Lizhen Qu, Iryna Gurevych
Abstract
Peer review is a cornerstone of quality control in scientific publishing. With the increasing workload, the unintended use of 'quick' heuristics, referred to as lazy thinking, has emerged as a recurring issue compromising review quality. Automated methods to detect such heuristics can help improve the peer-reviewing process. However, there is limited NLP research on this issue, and no real-world dataset exists to support the development of detection tools. This work introduces LAZYREVIEW, a dataset of peer-review sentences annotated with finegrained lazy thinking categories. Our analysis reveals that Large Language Models (LLMs) struggle to detect these instances in a zeroshot setting. However, instruction-based finetuning on our dataset significantly boosts performance by 10-20 performance points, highlighting the importance of high-quality training data. Furthermore, a controlled experiment demonstrates that reviews revised with lazy thinking feedback are more comprehensive and actionable than those written without such feedback. We will release our dataset and the enhanced guidelines that can be used to train junior reviewers in the community. 1 Heuristics Description Example review segments The results are not surprising Many findings seem obvious in retrospect, but this does not mean that the community is already aware of them and can use them as building blocks for future work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1ef579d2-78f4-4ee8-aa1d-e7e4eff0df02Cited by top-tier papers3
- Dancing in Chains: Strategic Persuasion in Academic Rebuttal via Theory of MindZhitao He, Zongwei Lyu, Yi R. FungICLR 2026 · 5 citations
- What Happens When Reviewers Receive AI Feedback in Their Reviews?Shiping Chen, Shu Zhong, Duncan P. Brumby, Anna L. CoxCHI 2026 · 1 citation
- The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for AuthorsAbdelrahman Sadallah, Tim Baumgärtner, Iryna Gurevych, Ted BriscoeEMNLP 2025
Builds on11
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- Exploring the Benefits of Training Expert Language Models over Instruction TuningJoel Jang, Seungone Kim, Seonghyeon Ye, Doyoung Kim et al.ICML 2023 · 97 citations
- Are Emergent Abilities in Large Language Models just In-Context Learning?Sheng Lu, Irina Bigoulaeva, Rachneet Sachdeva, Harish Tayyar Madabushi et al.ACL 2024 · 29 citations
- NLPeer: A Unified Resource for the Computational Study of Peer ReviewNils Dycke, Ilia Kuznetsov, Iryna GurevychACL 2023 · 17 citations
Related papers
- Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer ReviewSungduk Yu, Man Luo, Avinash Madasu, Vasudev Lal et al.ICLR 2026 · 24 citations
- Sem-Detect: Semantic Level Detection of AI Generated Peer-ReviewsAndré Duarte, Brian Tufts, Aditya Oke, Fei Fang et al.ICML 2026 · 1 citation
- This is not a Dataset: A Large Negation Benchmark to Challenge Large Language ModelsIker García-Ferrero, Begoña Altuna, Javier Álvez, Itziar Gonzalez-Dios et al.EMNLP 2023 · 8 citations
- CoCoNUTS: Concentrating on Content while Neglecting Uninformative Textual Styles for AI-Generated Peer Review DetectionYihan Chen, Jiawei Chen, Guozhao Mo, Xuanang Chen et al.ACL 2026 · 1 citation
- LLMs Assist NLP Researchers: Critique Paper (Meta-)ReviewingJiangshu Du, Yibo Wang, Wenting Zhao, Zhongfen Deng et al.EMNLP 2024 · 14 citations
