Mind the Blind Spots: A Focus-Level Evaluation Framework for LLM Reviews
Hyungyu Shin, Jingyu Tang, Yoonjoo Lee, Nayoung Kim, Hyunseung Lim, Ji Yong Cho, Hwajung Hong, Moontae Lee, Juho Kim
Abstract
Peer review underpins scientific progress, but it is increasingly strained by reviewer shortages and growing workloads. Large Language Models (LLMs) can automatically draft reviews now, but determining whether LLM-generated reviews are trustworthy requires systematic evaluation. Researchers have evaluated LLM reviews at either surface-level (e.g., BLEU and ROUGE) or content-level (e.g., specificity and factual accuracy). Yet it remains uncertain whether LLM-generated reviews attend to the same critical facets that human experts weighthe strengths and weaknesses that ultimately drive an accept-or-reject decision. We introduce a focus-level evaluation framework that operationalizes the focus as a normalized distribution of attention across predefined facets in paper reviews. Based on the framework, we developed an automatic focus-level evaluation pipeline based on two sets of facets: target (e.g., problem, method, and experiment) and aspect (e.g., validity, clarity, and novelty), leveraging 676 paper reviews 1 from OpenReview that consists of 3,657 strengths and weaknesses identified from human experts. The comparison of focus distributions between LLMs and human experts showed that the off-the-shelf LLMs consistently have a more biased focus towards examining technical validity while significantly overlooking novelty assessment when criticizing papers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Can AI Be a Good Peer Reviewer? A Survey of Peer Review Process, Evaluation, and the FutureSihong Wu, Owen Jiang, Yilun Zhao, Tiansheng Hu et al.ACL 2026 · 2 citations
- ReviewGrounder: Improving Review Substantiveness with Rubric-Guided, Tool-Integrated AgentsZhuofeng Li, Yi Lu, Dongfu Jiang, Haoxiang Zhang et al.ACL 2026 · 1 citation
- CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI ReviewersHexuan Deng, Xiaopeng Ke, Yichen Li, Ruina Hu et al.ICML 2026
Builds on5
- Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference AdjustmentRui Yang, Xiaoman Pan, Feng Luo, Shuang Qiu et al.ICML 2024 · 144 citations
- MetaWriter: Exploring the Potential and Perils of AI Writing Support in Scientific Peer ReviewLu Sun, Stone Tao, Junjie Hu, Steven P. DowCSCW 2024 · 36 citations
- Multi-Objective Preference Optimization: Improving Human Alignment of Generative ModelsAkhil Agnihotri, Rahul Jain, Deepak Ramachandran, Zheng WenICML 2026 · 15 citations
- LLMs Assist NLP Researchers: Critique Paper (Meta-)ReviewingJiangshu Du, Yibo Wang, Wenting Zhao, Zhongfen Deng et al.EMNLP 2024 · 14 citations
- The AI Review Lottery: Widespread AI-Assisted Peer Reviews Boost Paper Scores and Acceptance RatesGiuseppe Russo, Manoel Horta Ribeiro, Tim R. Davidson, Veniamin Veselovsky et al.CSCW 2025 · 10 citations
Related papers
- Auto-PRE: An Automatic and Cost-Efficient Peer-Review Framework for Language Generation EvaluationJunjie Chen, Weihang Su, Zhumin Chu, Haitao Li et al.AAAI 2026
- PiCO: Peer Review in LLMs based on Consistency OptimizationKun-Peng Ning, Shuo Yang, Yuyang Liu, Jia-Yu Yao et al.ICLR 2025
- AutoMedEval: Harnessing Language Models for Automatic Medical Capability EvaluationXiechi Zhang, Zetian Ouyang, Linlin Wang, Gerard de Melo et al.ACL 2025 · 1 citation
- LLMs can Perform Multi-Dimensional Analytic Writing Assessments: A Case Study of L2 Graduate-Level Academic English WritingZhengxiang Wang, Veronika Makarova, Zhi Li, Jordan Kodner et al.ACL 2025 · 5 citations
- Large Language Models for Automated Literature Review: An Evaluation of Reference Generation, Abstract Writing, and Review CompositionXuemei Tang, Xufeng Duan, Zhenguang G. CaiEMNLP 2025 · 5 citations
