NLPeer: A Unified Resource for the Computational Study of Peer Review
Nils Dycke, Ilia Kuznetsov, Iryna Gurevych
Abstract
Peer review constitutes a core component of scholarly publishing; yet it demands substantial expertise and training, and is susceptible to errors and biases. Various applications of NLP for peer reviewing assistance aim to support reviewers in this complex process, but the lack of clearly licensed datasets and multi-domain corpora prevent the systematic study of NLP for peer review. To remedy this, we introduce NLPEER -the first ethically sourced multidomain corpus of more than 5k papers and 11k review reports from five different venues. In addition to the new datasets of paper drafts, cameraready versions and peer reviews from the NLP community, we establish a unified data representation and augment previous peer review datasets to include parsed and structured paper representations, rich metadata and versioning information. We complement our resource with implementations and analysis of three reviewing assistance tasks, including a novel guided skimming task. Our work paves the path towards systematic, multi-faceted, evidencebased study of peer review in NLP and beyond. The data 1 and code 2 are publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer ReviewsWeixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp et al.ICML 2024 · 213 citations
- LLMs Assist NLP Researchers: Critique Paper (Meta-)ReviewingJiangshu Du, Yibo Wang, Wenting Zhao, Zhongfen Deng et al.EMNLP 2024 · 14 citations
- Help Me Write a Story: Evaluating LLMs' Ability to Generate Writing FeedbackHannah Rashkin, Elizabeth Clark, Fantine Huot, Mirella LapataACL 2025 · 7 citations
- Are Large Language Models Good Classifiers? A Study on Edit Intent Classification in Scientific Document RevisionsQian Ruan, Ilia Kuznetsov, Iryna GurevychEMNLP 2024 · 2 citations
- Closing the Loop: Learning to Generate Writing Feedback via Language Model Simulated Student RevisionsInderjeet Nair, Jiaye Tan, Xiaotian Su, Anne Gere et al.EMNLP 2024 · 2 citations
Builds on4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- On the Stability of Fine-tuning BERT: Misconceptions, Explanations, and Strong BaselinesMarius Mosbach, Maksym Andriushchenko, Dietrich KlakowICLR 2021 · 448 citations
- APE: Argument Pair Extraction from Peer Review and Rebuttal via Multi-task LearningLiying Cheng, Lidong Bing, Qian Yu, Wei Lu et al.EMNLP 2020 · 56 citations
- Prior and Prejudice: The Novice Reviewers' Bias against Resubmissions in Conference Peer ReviewIvan Stelmakh, Nihar B. Shah, Aarti Singh, Hal Daumé IIICSCW 2021 · 17 citations
Related papers
- Re3: A Holistic Framework and Dataset for Modeling Collaborative Document RevisionQian Ruan, Ilia Kuznetsov, Iryna GurevychACL 2024
- Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer ReviewSungduk Yu, Man Luo, Avinash Madasu, Vasudev Lal et al.ICLR 2026 · 24 citations
- Can AI Be a Good Peer Reviewer? A Survey of Peer Review Process, Evaluation, and the FutureSihong Wu, Owen Jiang, Yilun Zhao, Tiansheng Hu et al.ACL 2026 · 2 citations
- LazyReview: A Dataset for Uncovering Lazy Thinking in NLP Peer ReviewsSukannya Purkayastha, Zhuang Li, Anne Lauscher, Lizhen Qu et al.ACL 2025 · 1 citation
- Sem-Detect: Semantic Level Detection of AI Generated Peer-ReviewsAndré Duarte, Brian Tufts, Aditya Oke, Fei Fang et al.ICML 2026 · 1 citation
