Better than Average: Paired Evaluation of NLP systems
Maxime Peyrard, Wei Zhao, Steffen Eger, Robert West
Abstract
Evaluation in NLP is usually done by comparing the scores of competing systems independently averaged over a common set of test instances. In this work, we question the use of averages for aggregating evaluation scores into a final number used to decide which system is best, since the average, as well as alternatives such as the median, ignores the pairing arising from the fact that systems are evaluated on the same test instances. We illustrate the importance of taking the instancelevel pairing of evaluation scores into account and demonstrate, both theoretically and empirically, the advantages of aggregation methods based on pairwise comparisons, such as the Bradley-Terry (BT) model, a mechanism based on the estimated probability that a given system scores better than another on the test set. By re-evaluating 296 real NLP evaluation setups across four tasks and 18 evaluation metrics, we show that the choice of aggregation mechanism matters and yields different conclusions as to which systems are state of the art in about 30% of the setups. To facilitate the adoption of pairwise evaluation, we release a practical tool for performing the full analysis of evaluation scores with the mean, median, BT, and two variants of BT (Elo and TrueSkill), alongside functionality for appropriate statistical testing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3665d8ea-1ce4-4e34-88e6-9a924f5f0691Cited by top-tier papers10
- The MultiBERTs: BERT Reproductions for Robustness AnalysisThibault Sellam, Steve Yadlowsky, Ian Tenney, Jason Wei et al.ICLR 2022 · 106 citations
- What are the best Systems? New Perspectives on NLP BenchmarkingPierre Colombo, Nathan Noiry, Ekhine Irurozki, Stéphan ClémençonNeurIPS 2022 · 20 citations
- Invariant Language ModelingMaxime Peyrard, Sarvjeet Singh Ghotra, Martin Josifoski, Vidhan Agarwal et al.EMNLP 2022 · 8 citations
- Descartes: Generating Short Descriptions of Wikipedia ArticlesMarija Sakota, Maxime Peyrard, Robert WestWWW 2023 · 6 citations
- DecipherPref: Analyzing Influential Factors in Human Preference Judgments via GPT-4Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang et al.EMNLP 2023 · 3 citations
Builds on7
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Statistical Power and Translationese in Machine Translation EvaluationYvette Graham, Barry Haddow, Philipp KoehnEMNLP 2020 · 82 citations
- On the Limitations of Cross-lingual Encoders as Exposed by Reference-Free Machine Translation EvaluationWei Zhao, Goran Glavas, Maxime Peyrard, Yang Gao et al.ACL 2020 · 53 citations
- Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation MetricsNitika Mathur, Timothy Baldwin, Trevor CohnACL 2020 · 14 citations
- BLEU might be Guilty but References are not InnocentMarkus Freitag, David Grangier, Isaac CaswellEMNLP 2020 · 13 citations
Related papers
- Who can we trust? LLM-as-a-jury for Comparative AssessmentMengjie Qian, Guangzhi Sun, Mark Gales, Kate KnillICML 2026 · 5 citations
- Correct Looks Better: Pairwise Comparisons Reveal Accuracy RankingsMina Remeli, Moritz HardtICML 2026
- Elo Uncovered: Robustness and Best Practices in Language Model EvaluationMeriem Boubdir, Edward Kim, Beyza Ermis, Sara Hooker et al.NeurIPS 2024 · 94 citations
- Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie CalibrationDaniel Deutsch, George F. Foster, Markus FreitagEMNLP 2023 · 14 citations
- Rethinking Reward Modeling in Preference-based Large Language Model AlignmentHao Sun, Yunyi Shen, Jean-Francois TonICLR 2025
