Lune

CHI2026顶会

All Accept, No Reject: Evaluating LLMs as "Peer" Reviewers

Nitin Verma, Asheley R. Landrum

2026年份
2被引次数

摘要

An exponential rise in manuscript submission volume has strained the peer review system, prompting interest in automation from overburdened scholars and publishers. We systematically evaluated GPT-4.1, GPT-4o, o1, o3, o3-mini, and GPT-5 as “peer” reviewers, comparing their evaluation and acceptance of 137 manuscripts from an open dataset (PeerRead) with the corresponding human-generated reviews. While o3’s and GPT-5’s acceptance rates were close to the human benchmark (∼ 67% of submissions), others approved nearly every paper (>98%); all models performed extremely poorly on accuracy, precision, and recall metrics. To probe this striking “yes-bias”, we profiled the LLMs using Schwartz’s Portrait Values Questionnaire (PVQ-RR) and found that all LLMs emphasized self-transcendence and openness-to-change and de-emphasized conservation and self-enhancement. We argue that value orientations of LLMs we investigated are misaligned with the values underpinning peer review, and suggest new research on aligning AI judgment systems with human goals in this context.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖