Are LLM-based Evaluators Confusing NLG Quality Criteria?
Xinyu Hu, Mingqi Gao, Sen Hu, Yang Zhang, Yicheng Chen, Teng Xu, Xiaojun Wan
摘要
Some prior work has shown that LLMs perform well in NLG evaluation for different tasks. However, we discover that LLMs seem to confuse different evaluation criteria, which reduces their reliability. For further verification, we first consider avoiding issues of inconsistent conceptualization and vague expression in existing NLG quality criteria themselves. So we summarize a clear hierarchical classification system for 11 common aspects with corresponding different criteria from previous studies involved. Inspired by behavioral testing, we elaborately design 18 types of aspect-targeted perturbation attacks for fine-grained analysis of the evaluation behaviors of different LLMs. We also conduct human annotations beyond the guidance of the classification system to validate the impact of the perturbations. Our experimental results reveal confusion issues inherent in LLMs, as well as other noteworthy phenomena, and necessitate further research and improvements for LLM-based evaluation. * Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Beyond the Surface: Enhancing LLM-as-a-Judge Alignment with Human via Internal RepresentationsPeng Lai, Jianjie Zheng, Sijie Cheng, Yun Chen 等NeurIPS 2025 · 被引用 16 次
- FRABench and UFEval: Unified Fine-grained Evaluation with Task and Aspect GeneralizationShibo Hong, Jiahao Ying, Haiyuan Liang, Mengdi Zhang 等ICLR 2026 · 被引用 2 次
- Themis: A Reference-free NLG Evaluation Language Model with Flexibility and InterpretabilityXinyu Hu, Li Lin, Mingqi Gao, Xunjian Yin 等EMNLP 2024 · 被引用 2 次
- Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt EnsemblesEric Slyman, Md. Mehrab Tanjim, Kushal Kafle, Stefan LeeICCV 2025 · 被引用 1 次
- RevisEval: Improving LLM-as-a-Judge via Response-Adapted ReferencesQiyuan Zhang, Yufei Wang, Tiezheng Yu, Yuxin Jiang 等ICLR 2025
它引用的顶会 Paper12
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Can Large Language Models Be an Alternative to Human Evaluations?David Cheng-Han Chiang, Hung-yi LeeACL 2023 · 被引用 254 次
- Evaluating the Factual Consistency of Abstractive Text SummarizationWojciech Kryscinski, Bryan McCann, Caiming Xiong, Richard SocherEMNLP 2020 · 被引用 67 次
- A Comprehensive Analysis of the Effectiveness of Large Language Models as Automatic Dialogue EvaluatorsChen Zhang, Luis Fernando D'Haro, Yiming Chen, Malu Zhang 等AAAI 2024 · 被引用 57 次
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 被引用 51 次
相关 Paper
- A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better InterpretabilityXinyu Hu, Mingqi Gao, Li Lin, Zhenghan Yu 等ACL 2025
- Leveraging Large Language Models for NLG Evaluation: Advances and ChallengesZhen Li, Xiaohan Xu, Tao Shen, Can Xu 等EMNLP 2024 · 被引用 17 次
- CheckEval: A reliable LLM-as-a-Judge framework for evaluating text generation using checklistsYukyung Lee, JoongHoon Kim, Jaehee Kim, Hyowon Cho 等EMNLP 2025 · 被引用 2 次
- ConSiDERS-The-Human Evaluation Framework: Rethinking Human Evaluation for Generative Large Language ModelsAparna Elangovan, Ling Liu, Lei Xu, Sravan Babu Bodapati 等ACL 2024 · 被引用 19 次
- LLMEval: A Preliminary Study on How to Evaluate Large Language ModelsYue Zhang, Ming Zhang, Haipeng Yuan, Shichun Liu 等AAAI 2024 · 被引用 29 次
