What are the best Systems? New Perspectives on NLP Benchmarking
Pierre Colombo, Nathan Noiry, Ekhine Irurozki, Stéphan Clémençon
Abstract
In Machine Learning, a benchmark refers to an ensemble of datasets associated with one or multiple metrics together with a way to aggregate different systems performances. They are instrumental in (i) assessing the progress of new methods along different axes and (ii) selecting the best systems for practical use. This is particularly the case for NLP with the development of large pre-trained models (e.g. GPT, BERT) that are expected to generalize well on a variety of tasks. While the community mainly focused on developing new datasets and metrics, there has been little interest in the aggregation procedure, which is often reduced to a simple average over various performance measures. However, this procedure can be problematic when the metrics are on a different scale, which may lead to spurious conclusions. This paper proposes a new procedure to rank systems based on their performance across different tasks. Motivated by the social choice theory, the final system ordering is obtained through aggregating the rankings induced by each task and is theoretically grounded. We conduct extensive numerical experiments (on over 270k scores) to assess the soundness of our approach both on synthetic and real scores (e.g. GLUE, EXTREM, SEVAL, TAC, FLICKR). In particular, we show that our method yields different conclusions on state-of-the-art systems than the mean-aggregation procedure while being both more reliable and robust. How to aggregate performances? The multi-tasks setting has been investigated in recent works that provide benchmark of state-of-the-art models across a great variety of tasks [28, 62, 80, 90, 108] , sometimes with more than fifty [2, 84, 85, 94] . These papers provide tables of scores across the considered tasks, but the only non-qualitative way to compare systems consists in averaging the performances across tasks and then ranking systems according to their mean score values. This is, for instance, done with the GLUE benchmark [91] and its derivatives [92] . However, taking the mean is seriously flawed since the different metrics are usually not on the same scales and can even be unbounded [23, 102] . Even a pre-processing renormalization scheme would fail to capture the intrinsic difficulty of the tasks. Contribution 1. Our first contribution is to provide a reliable tool to rank systems in a multi-tasks setting. We rely on a ranking aggregation procedure which, from a set of rankings induced by each criterion, returns a single ranking that somehow aggregates the former. This procedure, called the Kemeny consensus [52], can be seen as a voting rule and stems from the social choice theory [66] . Aggregation when instance-level information is available. As illustrated by Ruder [83], Zhong et al. [109], a fine-grained understanding of the model performance should include instance-level scores. If taking the mean is quite natural in the classification setting, this is not always the case, as recently pointed out by [73] in the NLG setting. In this article, the authors investigate pairwise comparison of NLG systems for a single metric (e.g. BLEU [71], ROUGE [59], METEOR [5, 35, 49] , CHRF [76, 77] , BertScore [105] ). They prove that a comparison based on the mean or the median of the scores across test utterances can be highly flawed. They rather advise to rely on the Bradley-Terry [10] pairwise comparison method, which consists, for two systems A and B, in computing the proportion of utterances on which A achieves a better score than B. Their work is a significant advance but remains limited to pairwise comparisons. Contribution 2. Our second contribution consists in going one step further than [73] by applying our ranking procedure to an arbitrarily large set of NLG systems with respect to a group of fixed criterion. Our evaluation methodology can be seen as a natural extension of [73] since it coincides with the latter in the particular case of pairwise comparison. In a more realistic multi-criteria scenario, we combine our two contributions and develop a two-stages ranking aggregation procedure which first aggregates along utterances and then along criteria. Experiments. Our two contributions rely on our aggregation procedure which is proved to be effective through several experiments. 1. We explain on a simple synthetic example the superiority of our approach compared to the mean-aggregation procedure and the pairwise-aggregation procedure, both in terms of consistency and robustness. 2. We use our ranking procedure on 10 multi-tasks / multi-criteria benchmarks and observe it leads to different conclusions than mean-and pairwise-aggregation procedures. 3. We argue our procedure is more robust by investigating its stability with respect to the addition of criteria and with respect to the addition of systems. Our code and the collected data will be released to accelerate the adoption of what we think is a reliable evaluation method for multi-tasks and multi-criteria benchmarks. 2 Problem Formula
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 750c9449-5a35-4e4d-b176-c2a6f09acfe7Cited by top-tier papers11
- Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic EncodersYupeng Hou, Jiacheng Li, Xiangjun Fu, Zhankui He et al.ACL 2026 · 346 citations
- InfoLM: A New Metric to Evaluate Summarization & Data2Text GenerationPierre Jean A. Colombo, Chloé Clavel, Pablo PiantanidaAAAI 2022 · 52 citations
- ClusterLLM: Large Language Models as a Guide for Text ClusteringYuwei Zhang, Zihan Wang, Jingbo ShangEMNLP 2023 · 43 citations
- Beyond Mahalanobis Distance for Textual OOD DetectionPierre Colombo, Eduardo Dadalto Câmara Gomes, Guillaume Staerman, Nathan Noiry et al.NeurIPS 2022 · 24 citations
- Inherent Trade-Offs between Diversity and Stability in Multi-Task BenchmarksGuanhua Zhang, Moritz HardtICML 2024 · 22 citations
Builds on16
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralisationJunjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig et al.ICML 2020 · 1,132 citations
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen et al.ICLR 2021 · 881 citations
- ExT5: Towards Extreme Multi-Task Scaling for Transfer LearningVamsi Aribandi, Yi Tay, Tal Schuster, Jinfeng Rao et al.ICLR 2022 · 237 citations
Related papers
- Better than Average: Paired Evaluation of NLP systemsMaxime Peyrard, Wei Zhao, Steffen Eger, Robert WestACL 2021
- SkillAggregation: Reference-free LLM-Dependent AggregationGuangzhi Sun, Anmol Kagrecha, Potsawee Manakul, Philip C. Woodland et al.ACL 2025 · 5 citations
- In Benchmarks We Trust ... Or Not?Ine Gevers, Victor De Marez, Jens Van Nooten, Jens Lemmens et al.EMNLP 2025 · 1 citation
- Robust AI Evaluation through Maximal LotteriesHadi Khalaf, Serena Wang, Daniel Halpern, Itai Shapira et al.ICML 2026 · 2 citations
- Can Machines Read Coding Manuals Yet? - A Benchmark for Building Better Language Models for Code UnderstandingIbrahim Abdelaziz, Julian Dolby, Jamie P. McCusker, Kavitha SrinivasAAAI 2022 · 7 citations
