How Robust are Model Rankings : A Leaderboard Customization Approach for Equitable Evaluation
Swaroop Mishra, Anjana Arunkumar
Abstract
Models that top leaderboards often perform unsatisfactorily when deployed in real world applications; this has necessitated rigorous and expensive pre-deployment model testing. A hitherto unexplored facet of model performance is: Are our leaderboards doing equitable evaluation? In this paper, we introduce a task-agnostic method to probe leaderboards by weighting samples based on their 'difficulty' level. We find that leaderboards can be adversarially attacked and top performing models may not always be the best models. We subsequently propose alternate evaluation metrics. Our experiments on 10 models show changes in model ranking and an overall reduction in previously reported performance- thus rectifying the overestimation of AI systems' capabilities. Inspired by behavioral testing principles, we further develop a prototype of a visual analytics tool that enables leaderboard revamping through customization, based on an end user's focus area. This helps users analyze models' strengths and weaknesses, and guides them in the selection of a model best suited for their application scenario. In a user study, members of various commercial product development teams, covering 5 focus areas, find that our prototype reduces pre-deployment development and testing effort by 41% on average.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0449a7ce-b4e5-4814-8326-023c83e54f46Cited by top-tier papers6
- Inherent Trade-Offs between Diversity and Stability in Multi-Task BenchmarksGuanhua Zhang, Moritz HardtICML 2024 · 22 citations
- ILDAE: Instance-Level Difficulty Analysis of Evaluation DataNeeraj Varshney, Swaroop Mishra, Chitta BaralACL 2022 · 21 citations
- PMU Tracker: A Visualization Platform for Epicentric Event Propagation Analysis in the Power GridAnjana Arunkumar, Andrea Pinceti, Lalitha Sankar, Chris BryanIEEE VIS 2022 · 9 citations
- ECBD: Evidence-Centered Benchmark Design for NLPYu Lu Liu, Su Lin Blodgett, Jackie C. K. Cheung, Vera Liao et al.ACL 2024 · 3 citations
- Metritocracy: Representative Metrics for Lite BenchmarksAriel D. Procaccia, Ben Schiffer, Serena Wang, Shirley ZhangNeurIPS 2025 · 2 citations
Builds on6
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- Adversarial Filters of Dataset BiasesRonan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers et al.ICML 2020 · 242 citations
- Selective Question Answering under Domain ShiftAmita Kamath, Robin Jia, Percy LiangACL 2020 · 121 citations
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 51 citations
Related papers
- Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?Pedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor et al.ACL 2021
- Exploiting Leaderboards for Large-Scale Distribution of Malicious ModelsAnshuman Suri, Harsh Chaudhari, Yuefeng Peng, Ali Naseh et al.S&P 2026
- When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model LeaderboardsNorah A. Alzahrani, Hisham Abdullah Alyahya, Yazeed Alnumay, Sultan Alrashed et al.ACL 2024 · 13 citations
- Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent EvaluationSayash Kapoor, Benedikt Stroebl, Peter Kirgis, Nitya Nadgir et al.ICLR 2026 · 86 citations
- Prompt-to-Leaderboard: Prompt-Adaptive LLM EvaluationsEvan Frick, Connor Chen, Joseph Tennyson, Tianle Li et al.ICML 2025
