LLMs Assist NLP Researchers: Critique Paper (Meta-)Reviewing
Jiangshu Du, Yibo Wang, Wenting Zhao, Zhongfen Deng, Shuaiqi Liu, Renze Lou, Henry Peng Zou, Pranav Narayanan Venkit, Nan Zhang, Mukund Srinath, Haoran Zhang, Vipul Gupta
Abstract
Claim: This work is not advocating the use of LLMs for paper (meta-)reviewing. Instead, we present a comparative analysis to identify and distinguish LLM activities from human activities. Two research goals: i) Enable better recognition of instances when someone implicitly uses LLMs for reviewing activities; ii) Increase community awareness that LLMs, and AI in general, are currently inadequate for performing tasks that require a high level of expertise and nuanced judgment. This work is motivated by two key trends. On one hand, large language models (LLMs) have shown remarkable versatility in various generative tasks such as writing, drawing, and question answering, significantly reducing the time required for many routine tasks. On the other hand, researchers, whose work is not only time-consuming but also highly expertisedemanding, face increasing challenges as they have to spend more time reading, writing, and reviewing papers. This raises the question: how can LLMs potentially assist researchers in alleviating their heavy workload? This study focuses on the topic of LLMs Assist NLP Researchers, particularly examining the effectiveness of LLM in assisting paper (meta-)reviewing and its recognizability. To address this, we constructed the ReviewCritique dataset, which includes two types of information: (i) NLP papers (initial submissions rather than camera-ready) with both human-written and LLM-generated reviews, and (ii) each review comes with "deficiency" labels and corresponding explanations for individual segments, annotated by experts. Using ReviewCritique, this study explores two threads of research questions: (i) "LLMs as Reviewers", how do reviews generated by LLMs compare with those written by humans in terms of quality and distinguishability? (ii) "LLMs as Metareviewers", how effectively can LLMs identify potential issues, such as Deficient or unprofessional review segments, within individual paper reviews? To our knowledge, this is the first work to provide such a comprehensive analysis. Our dataset is available at https://github.com/jiangshdd/ReviewCritique .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking ProcessMinjun Zhu, Yixuan Weng, Linyi Yang, Yue ZhangACL 2025 · 70 citations
- AgentReview: Exploring Peer Review Dynamics with LLM AgentsYiqiao Jin, Qinlin Zhao, Yiyang Wang, Hao Chen et al.EMNLP 2024 · 24 citations
- Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning AbilitiesJiayi Kuang, Haojing Huang, Yinghui Li, Xinnian Liang et al.NeurIPS 2025 · 11 citations
- RExBench: Can coding agents autonomously implement AI research extensions?Nicholas Edwards, Yukyung Lee, Yujun Audrey Mao, Yulu Qin et al.ACL 2026 · 10 citations
- Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMsJungsoo Park, Junmo Kang, Gabriel Stanovsky, Alan RitterACL 2025 · 4 citations
Builds on4
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- APE: Argument Pair Extraction from Peer Review and Rebuttal via Multi-task LearningLiying Cheng, Lidong Bing, Qian Yu, Wei Lu et al.EMNLP 2020 · 56 citations
- NLPeer: A Unified Resource for the Computational Study of Peer ReviewNils Dycke, Ilia Kuznetsov, Iryna GurevychACL 2023 · 17 citations
- A Dataset of Argumentative Dialogues on Scientific PapersFederico Ruggeri, Mohsen Mesgar, Iryna GurevychACL 2023
Related papers
- AAAR-1.0: Assessing AI's Potential to Assist ResearchRenze Lou, Hanzi Xu, Sijia Wang, Jiangshu Du et al.ICML 2025
- CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model GenerationPei Ke, Bosi Wen, Andrew Feng, Xiao Liu et al.ACL 2024 · 9 citations
- Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer ReviewSungduk Yu, Man Luo, Avinash Madasu, Vasudev Lal et al.ICLR 2026 · 24 citations
- Large Language Models for Automated Literature Review: An Evaluation of Reference Generation, Abstract Writing, and Review CompositionXuemei Tang, Xufeng Duan, Zhenguang G. CaiEMNLP 2025 · 5 citations
- CriticEval: Evaluating Large-scale Language Model as CriticTian Lan, Wenwei Zhang, Chen Xu, Heyan Huang et al.NeurIPS 2024 · 26 citations
