The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors
Abdelrahman Sadallah, Tim Baumgärtner, Iryna Gurevych, Ted Briscoe
Abstract
Providing constructive feedback to paper authors is a core component of peer review. With reviewers increasingly having less time to perform reviews, automated support systems are required to ensure high reviewing quality, thus making the feedback in reviews useful for authors. To this end, we identify four key aspects of review comments (individual points in weakness sections of reviews) that drive the utility for authors: Actionability, Grounding & Specificity, Verifiability, and Helpfulness. To enable evaluation and development of models assessing review comments, we introduce the RevUtil dataset. We collect 1,430 human-labeled review comments and scale our data with 10k synthetically labeled comments for training purposes. The synthetic data additionally contains rationales, i.e., explanations for the aspect score of a review comment. Employing the RevUtil dataset, we benchmark fine-tuned models for assessing review comments on these aspects and generating rationales. Our experiments demonstrate that these fine-tuned models achieve agreement levels with humans comparable to, and in some cases exceeding, those of powerful closed models like GPT-4o. Our analysis further reveals that machinegenerated reviews generally underperform human reviews on our four aspects. 1 Review Utility Evaluation Peer Review "The number of baselines is a bit small, which degrades its universality and generality." Aspect: Actionability Score: 2/5 Rationale: "The review [...] does not provide specific guidance [...] such as suggesting additional baselines to include or explaining how to enhance the universality and generality. The action is implicit, as the authors need to infer that they should add more baselines, and it is vague because it lacks concrete steps for improvement. [...]" Extract individual review comments from peer reviews Aspects annotation for each comment on 1-5 Scale Aggregate human annotations into full, major and low agreement Scale annotations with GPT-4o aligning with human data and generating score rationale Full (3/3) "The points raised in Section 5 would benefit from more in-depth analysis" Aspect Definitions Aspect Annotation Training & Inference De!ne RevUtil Aspects based on Aspect Literature and Reviewer Guidelines Actionability -Degree to which a comment explicitly states a concrete action to perform to improve the contribution Grounding + Speci!city -Explicit link to a part of the paper and speci!c details what to improve in this part Veri!ability -Measures to which extend the comment provides evidence and rationales to support its claims Helpfulness -Overall judgment of the review comment on how helpful it is for an author to improve their work Aspect Lit. Reviewer Guidelines Peer Review Low Train small-scale, practical models on to predict aspect scores and generate rationales Evaluate models on human and synthetic data RevUtil Synthetic 3 5 2 2 4 3 4 3 5 3 4 4 Comment
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa7621d5-bf32-4246-bb47-62e326183635Cited by top-tier papers3
- Can AI Be a Good Peer Reviewer? A Survey of Peer Review Process, Evaluation, and the FutureSihong Wu, Owen Jiang, Yilun Zhao, Tiansheng Hu et al.ACL 2026 · 2 citations
- Reward Modeling for Scientific Writing EvaluationFurkan Sahinuç, Subhabrata Dutta, Iryna GurevychACL 2026 · 2 citations
- CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI ReviewersHexuan Deng, Xiaopeng Ke, Yichen Li, Ruina Hu et al.ICML 2026
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
Related papers
- ARIES: A Corpus of Scientific Paper Edits Made in Response to Peer ReviewsMike D'Arcy, Alexis Ross, Erin Bransom, Bailey Kuehl et al.ACL 2024 · 3 citations
- ReviewGrounder: Improving Review Substantiveness with Rubric-Guided, Tool-Integrated AgentsZhuofeng Li, Yi Lu, Dongfu Jiang, Haoxiang Zhang et al.ACL 2026 · 1 citation
- ReviewRL: Towards Automated Scientific Review with RLSihang Zeng, Kai Tian, Kaiyan Zhang, Yuru Wang et al.EMNLP 2025
- LazyReview: A Dataset for Uncovering Lazy Thinking in NLP Peer ReviewsSukannya Purkayastha, Zhuang Li, Anne Lauscher, Lizhen Qu et al.ACL 2025 · 1 citation
- Agent Reviewers: Domain-specific Multimodal Agents with Shared Memory for Paper ReviewKai Lu, Shixiong Xu, Jinqiu Li, Kun Ding et al.ICML 2025
