CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI Reviewers
Hexuan Deng, Xiaopeng Ke, Yichen Li, Ruina Hu, Dehao Huang, Derek F. Wong, Yue Wang, Xuebo Liu, Min zhang
摘要
Despite the rapid development of AI reviewers, evaluating such systems remains challenging: metrics favor overlap with human reviews over correctness. However, since human reviews often cover only a subset of salient issues and sometimes contain mistakes, they are unreliable as gold references. To address this, we build categoryspecific benchmark subsets and skip evaluation when the corresponding human reviews are missing to strengthen Completeness. We also leverage reviewer-author-meta-review discussions as expert annotations and filter unreliable reviews accordingly to strengthen Correctness. Finally, we introduce CoCoReviewBench, which curates 3,900 papers from ICLR and NeurIPS to enable reliable and fine-grained evaluation of AI reviewers. Analysis shows that AI reviewers remain limited in correctness and are prone to hallucinations, and highlights reasoning models as more effective reviewers, motivating further directions for improving AI reviewers. Benchmarks and models are available at https://github.c om/hexuandeng/CoCoReviewBench .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- AgentReview: Exploring Peer Review Dynamics with LLM AgentsYiqiao Jin, Qinlin Zhao, Yiyang Wang, Hao Chen 等EMNLP 2024 · 被引用 24 次
- Improving Simultaneous Machine Translation with Monolingual DataHexuan Deng, Liang Ding, Xuebo Liu, Meishan Zhang 等AAAI 2023 · 被引用 19 次
- Mind the Blind Spots: A Focus-Level Evaluation Framework for LLM ReviewsHyungyu Shin, Jingyu Tang, Yoonjoo Lee, Nayoung Kim 等EMNLP 2025 · 被引用 2 次
- Benchmarking LLMs' Judgments with No Gold StandardShengwei Xu, Yuxuan Lu, Grant Schoenebeck, Yuqing KongICLR 2025
相关 Paper
- RATE: Reviewer Profiling and Annotation-free Training for Expertise Ranking in Peer Review SystemsWeicong Liu, Zixuan Yang, Yibo Zhao, Xiang LiACL 2026
- ReviewGrounder: Improving Review Substantiveness with Rubric-Guided, Tool-Integrated AgentsZhuofeng Li, Yi Lu, Dongfu Jiang, Haoxiang Zhang 等ACL 2026 · 被引用 1 次
- RubricBench: Aligning Model-Generated Rubrics with Human StandardsJunyi Zhou, Qiyuan Zhang, Yufei Wang, Fuyuan Lyu 等ACL 2026 · 被引用 7 次
- CogBench: a large language model walks into a psychology labJulian Coda-Forno, Marcel Binz, Jane X. Wang, Eric SchulzICML 2024 · 被引用 60 次
- CoCoNUTS: Concentrating on Content while Neglecting Uninformative Textual Styles for AI-Generated Peer Review DetectionYihan Chen, Jiawei Chen, Guozhao Mo, Xuanang Chen 等ACL 2026 · 被引用 1 次
