The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors
Abdelrahman Sadallah, Tim Baumgärtner, Iryna Gurevych, Ted Briscoe
摘要
Providing constructive feedback to paper authors is a core component of peer review. With reviewers increasingly having less time to perform reviews, automated support systems are required to ensure high reviewing quality, thus making the feedback in reviews useful for authors. To this end, we identify four key aspects of review comments (individual points in weakness sections of reviews) that drive the utility for authors: Actionability, Grounding & Specificity, Verifiability, and Helpfulness. To enable evaluation and development of models assessing review comments, we introduce the RevUtil dataset. We collect 1,430 human-labeled review comments and scale our data with 10k synthetically labeled comments for training purposes. The synthetic data additionally contains rationales, i.e., explanations for the aspect score of a review comment. Employing the RevUtil dataset, we benchmark fine-tuned models for assessing review comments on these aspects and generating rationales. Our experiments demonstrate that these fine-tuned models achieve agreement levels with humans comparable to, and in some cases exceeding, those of powerful closed models like GPT-4o. Our analysis further reveals that machinegenerated reviews generally underperform human reviews on our four aspects. 1 Review Utility Evaluation Peer Review "The number of baselines is a bit small, which degrades its universality and generality." Aspect: Actionability Score: 2/5 Rationale: "The review [...] does not provide specific guidance [...] such as suggesting additional baselines to include or explaining how to enhance the universality and generality. The action is implicit, as the authors need to infer that they should add more baselines, and it is vague because it lacks concrete steps for improvement. [...]" Extract individual review comments from peer reviews Aspects annotation for each comment on 1-5 Scale Aggregate human annotations into full, major and low agreement Scale annotations with GPT-4o aligning with human data and generating score rationale Full (3/3) "The points raised in Section 5 would benefit from more in-depth analysis" Aspect Definitions Aspect Annotation Training & Inference De!ne RevUtil Aspects based on Aspect Literature and Reviewer Guidelines Actionability -Degree to which a comment explicitly states a concrete action to perform to improve the contribution Grounding + Speci!city -Explicit link to a part of the paper and speci!c details what to improve in this part Veri!ability -Measures to which extend the comment provides evidence and rationales to support its claims Helpfulness -Overall judgment of the review comment on how helpful it is for an author to improve their work Aspect Lit. Reviewer Guidelines Peer Review Low Train small-scale, practical models on to predict aspect scores and generate rationales Evaluate models on human and synthetic data RevUtil Synthetic 3 5 2 2 4 3 4 3 5 3 4 4 Comment
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Can AI Be a Good Peer Reviewer? A Survey of Peer Review Process, Evaluation, and the FutureSihong Wu, Owen Jiang, Yilun Zhao, Tiansheng Hu 等ACL 2026 · 被引用 2 次
- Reward Modeling for Scientific Writing EvaluationFurkan Sahinuç, Subhabrata Dutta, Iryna GurevychACL 2026 · 被引用 2 次
- CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI ReviewersHexuan Deng, Xiaopeng Ke, Yichen Li, Ruina Hu 等ICML 2026
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
相关 Paper
- ARIES: A Corpus of Scientific Paper Edits Made in Response to Peer ReviewsMike D'Arcy, Alexis Ross, Erin Bransom, Bailey Kuehl 等ACL 2024 · 被引用 3 次
- ReviewGrounder: Improving Review Substantiveness with Rubric-Guided, Tool-Integrated AgentsZhuofeng Li, Yi Lu, Dongfu Jiang, Haoxiang Zhang 等ACL 2026 · 被引用 1 次
- ReviewRL: Towards Automated Scientific Review with RLSihang Zeng, Kai Tian, Kaiyan Zhang, Yuru Wang 等EMNLP 2025
- LazyReview: A Dataset for Uncovering Lazy Thinking in NLP Peer ReviewsSukannya Purkayastha, Zhuang Li, Anne Lauscher, Lizhen Qu 等ACL 2025 · 被引用 1 次
- Agent Reviewers: Domain-specific Multimodal Agents with Shared Memory for Paper ReviewKai Lu, Shixiong Xu, Jinqiu Li, Kun Ding 等ICML 2025
