Active Evaluation: Efficient NLG Evaluation with Few Pairwise Comparisons
Akash Kumar Mohankumar, Mitesh M. Khapra
Abstract
Recent studies have shown the advantages of evaluating NLG systems using pairwise comparisons as opposed to direct assessment. Given k systems, a naive approach for identifying the top-ranked system would be to uniformly obtain pairwise comparisons from all k 2 pairs of systems. However, this can be very expensive as the number of human annotations required would grow quadratically with k. In this work, we introduce Active Evaluation, a framework to efficiently identify the top-ranked system by actively choosing system pairs for comparison using dueling bandit algorithms. We perform extensive experiments with 13 dueling bandits algorithms on 13 NLG evaluation datasets spanning 5 tasks and show that the number of human annotations can be reduced by 80%. To further reduce the number of human annotations, we propose model-based dueling bandit algorithms which combine automatic evaluation metrics with human evaluations. Specifically, we eliminate sub-optimal systems even before the human annotation process and perform human evaluations only on test examples where the automatic metric is highly uncertain. This reduces the number of human annotations required further by 89%. In effect, we show that identifying the top-ranked system requires only a few hundred human annotations, which grow linearly with k. Lastly, we provide practical recommendations and best practices to identify the top-ranked system efficiently. Our code has been made publicly available at https://github.com/akashkm99/duelnlg
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- A Survey of Active Learning for Natural Language ProcessingZhisong Zhang, Emma Strubell, Eduard H. HovyEMNLP 2022 · 60 citations
- Better than Random: Reliable NLG Human Evaluation with Constrained Active SamplingJie Ruan, Xiao Pu, Mingqi Gao, Xiaojun Wan et al.AAAI 2024 · 8 citations
- Label-Efficient Model Selection for Text GenerationShir Ashury-Tahan, Ariel Gera, Benjamin Sznajder, Leshem Choshen et al.ACL 2024 · 1 citation
Builds on4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
- Beyond User Self-Reported Likert Scale Ratings: A Comparison Model for Automatic Dialog EvaluationWeixin Liang, James Zou, Zhou YuACL 2020 · 25 citations
Related papers
- CUPID in the Model Zoo: Online Matchmaking for Selecting Your Dream LLMSon Nguyen, Xinyuan Liu, Ransalu SenanayakeICML 2026
- Human Preferences as Dueling BanditsXinyi Yan, Chengxi Luo, Charles L. A. Clarke, Nick Craswell et al.SIGIR 2022 · 9 citations
- Efficient LLM Comparative Assessment: A Product of Experts Framework for Pairwise ComparisonsAdian Liusie, Vatsal Raina, Yassir Fathullah, Mark J. F. GalesEMNLP 2024 · 5 citations
- A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better InterpretabilityXinyu Hu, Mingqi Gao, Li Lin, Zhenghan Yu et al.ACL 2025
- Better than Average: Paired Evaluation of NLP systemsMaxime Peyrard, Wei Zhao, Steffen Eger, Robert WestACL 2021
