Lune

CHI2026Top-tier venue

All Accept, No Reject: Evaluating LLMs as "Peer" Reviewers

Nitin Verma, Asheley R. Landrum

2026Year
2Citations

Abstract

An exponential rise in manuscript submission volume has strained the peer review system, prompting interest in automation from overburdened scholars and publishers. We systematically evaluated GPT-4.1, GPT-4o, o1, o3, o3-mini, and GPT-5 as “peer” reviewers, comparing their evaluation and acceptance of 137 manuscripts from an open dataset (PeerRead) with the corresponding human-generated reviews. While o3’s and GPT-5’s acceptance rates were close to the human benchmark (∼ 67% of submissions), others approved nearly every paper (>98%); all models performed extremely poorly on accuracy, precision, and recall metrics. To probe this striking “yes-bias”, we profiled the LLMs using Schwartz’s Portrait Values Questionnaire (PVQ-RR) and found that all LLMs emphasized self-transcendence and openness-to-change and de-emphasized conservation and self-enhancement. We argue that value orientations of LLMs we investigated are misaligned with the values underpinning peer review, and suggest new research on aligning AI judgment systems with human goals in this context.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines