LazyReview: A Dataset for Uncovering Lazy Thinking in NLP Peer Reviews
Sukannya Purkayastha, Zhuang Li, Anne Lauscher, Lizhen Qu, Iryna Gurevych
摘要
Peer review is a cornerstone of quality control in scientific publishing. With the increasing workload, the unintended use of 'quick' heuristics, referred to as lazy thinking, has emerged as a recurring issue compromising review quality. Automated methods to detect such heuristics can help improve the peer-reviewing process. However, there is limited NLP research on this issue, and no real-world dataset exists to support the development of detection tools. This work introduces LAZYREVIEW, a dataset of peer-review sentences annotated with finegrained lazy thinking categories. Our analysis reveals that Large Language Models (LLMs) struggle to detect these instances in a zeroshot setting. However, instruction-based finetuning on our dataset significantly boosts performance by 10-20 performance points, highlighting the importance of high-quality training data. Furthermore, a controlled experiment demonstrates that reviews revised with lazy thinking feedback are more comprehensive and actionable than those written without such feedback. We will release our dataset and the enhanced guidelines that can be used to train junior reviewers in the community. 1 Heuristics Description Example review segments The results are not surprising Many findings seem obvious in retrospect, but this does not mean that the community is already aware of them and can use them as building blocks for future work.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Dancing in Chains: Strategic Persuasion in Academic Rebuttal via Theory of MindZhitao He, Zongwei Lyu, Yi R. FungICLR 2026 · 被引用 5 次
- What Happens When Reviewers Receive AI Feedback in Their Reviews?Shiping Chen, Shu Zhong, Duncan P. Brumby, Anna L. CoxCHI 2026 · 被引用 1 次
- The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for AuthorsAbdelrahman Sadallah, Tim Baumgärtner, Iryna Gurevych, Ted BriscoeEMNLP 2025
它引用的顶会 Paper11
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- Exploring the Benefits of Training Expert Language Models over Instruction TuningJoel Jang, Seungone Kim, Seonghyeon Ye, Doyoung Kim 等ICML 2023 · 被引用 97 次
- Are Emergent Abilities in Large Language Models just In-Context Learning?Sheng Lu, Irina Bigoulaeva, Rachneet Sachdeva, Harish Tayyar Madabushi 等ACL 2024 · 被引用 29 次
- NLPeer: A Unified Resource for the Computational Study of Peer ReviewNils Dycke, Ilia Kuznetsov, Iryna GurevychACL 2023 · 被引用 17 次
相关 Paper
- Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer ReviewSungduk Yu, Man Luo, Avinash Madasu, Vasudev Lal 等ICLR 2026 · 被引用 24 次
- Sem-Detect: Semantic Level Detection of AI Generated Peer-ReviewsAndré Duarte, Brian Tufts, Aditya Oke, Fei Fang 等ICML 2026 · 被引用 1 次
- This is not a Dataset: A Large Negation Benchmark to Challenge Large Language ModelsIker García-Ferrero, Begoña Altuna, Javier Álvez, Itziar Gonzalez-Dios 等EMNLP 2023 · 被引用 8 次
- CoCoNUTS: Concentrating on Content while Neglecting Uninformative Textual Styles for AI-Generated Peer Review DetectionYihan Chen, Jiawei Chen, Guozhao Mo, Xuanang Chen 等ACL 2026 · 被引用 1 次
- LLMs Assist NLP Researchers: Critique Paper (Meta-)ReviewingJiangshu Du, Yibo Wang, Wenting Zhao, Zhongfen Deng 等EMNLP 2024 · 被引用 14 次
