BatchEval: Towards Human-like Text Evaluation
Peiwen Yuan, Shaoxiong Feng, Yiwei Li, Xinglin Wang, Boyuan Pan, Heda Wang, Yao Hu, Kan Li
摘要
Significant progress has been made in automatic text evaluation with the introduction of large language models (LLMs) as evaluators. However, current sample-wise evaluation paradigm suffers from the following issues: (1) Sensitive to prompt design; (2) Poor resistance to noise; (3) Inferior ensemble performance with static reference. Inspired by the fact that humans treat both criterion definition and inter sample comparison as references for evaluation, we propose BATCHEVAL, a paradigm that conducts batch-wise evaluation iteratively to alleviate the above problems. We explore variants under this paradigm and confirm the optimal settings are two stage procedure with heterogeneous batch composition strategy and decimal scoring format. Comprehensive experiments across 3 LLMs on 4 text evaluation tasks demonstrate that BATCHEVAL outperforms state-of-the-art methods by 10.5% on Pearson correlations with only 64% API cost on average. Further analyses have been conducted to verify the robustness, generalization, and working mechanism of BATCHEVAL 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software EngineeringRuiqi Wang, Jiyu Guo, Cuiyun Gao, Guodong Fan 等ISSTA 2025 · 被引用 27 次
- Large Language Model as an Assignment Evaluator: Insights, Feedback, and Challenges in a 1000+ Student CourseCheng-Han Chiang, Wei-Chih Chen, Chun-Yi Kuan, Chienchou Yang 等EMNLP 2024 · 被引用 14 次
- How to Engage your Readers? Generating Guiding Questions to Promote Active ReadingPeng Cui, Vilém Zouhar, Xiaoyu Zhang, Mrinmaya SachanACL 2024 · 被引用 2 次
- EmoHarbor: Evaluating Personalized Emotional Support by Simulating the User's Internal WorldJing Ye, Lu Xiang, Yaping Zhang, Chengqing ZongACL 2026 · 被引用 2 次
- UniCBE: An Uniformity-driven Comparing Based Evaluation Framework with Unified Multi-Objective OptimizationPeiwen Yuan, Shaoxiong Feng, Yiwei Li, Xinglin Wang 等ICLR 2025
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang 等NeurIPS 2023 · 被引用 948 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
相关 Paper
- RevisEval: Improving LLM-as-a-Judge via Response-Adapted ReferencesQiyuan Zhang, Yufei Wang, Tiezheng Yu, Yuxin Jiang 等ICLR 2025
- RepEval: Effective Text Evaluation with LLM RepresentationShuqian Sheng, Yi Xu, Tianhang Zhang, Zanwei Shen 等EMNLP 2024 · 被引用 5 次
- Themis: A Reference-free NLG Evaluation Language Model with Flexibility and InterpretabilityXinyu Hu, Li Lin, Mingqi Gao, Xunjian Yin 等EMNLP 2024 · 被引用 2 次
- HypoEval: Hypothesis-Guided Evaluation for Natural Language GenerationMingxuan Li, Hanchen Li, Chenhao TanACL 2026 · 被引用 1 次
- Systematic Task Exploration with LLMs: A Study in Citation Text GenerationFurkan Sahinuç, Ilia Kuznetsov, Yufang Hou, Iryna GurevychACL 2024
