How Robust are Model Rankings : A Leaderboard Customization Approach for Equitable Evaluation
Swaroop Mishra, Anjana Arunkumar
摘要
Models that top leaderboards often perform unsatisfactorily when deployed in real world applications; this has necessitated rigorous and expensive pre-deployment model testing. A hitherto unexplored facet of model performance is: Are our leaderboards doing equitable evaluation? In this paper, we introduce a task-agnostic method to probe leaderboards by weighting samples based on their 'difficulty' level. We find that leaderboards can be adversarially attacked and top performing models may not always be the best models. We subsequently propose alternate evaluation metrics. Our experiments on 10 models show changes in model ranking and an overall reduction in previously reported performance- thus rectifying the overestimation of AI systems' capabilities. Inspired by behavioral testing principles, we further develop a prototype of a visual analytics tool that enables leaderboard revamping through customization, based on an end user's focus area. This helps users analyze models' strengths and weaknesses, and guides them in the selection of a model best suited for their application scenario. In a user study, members of various commercial product development teams, covering 5 focus areas, find that our prototype reduces pre-deployment development and testing effort by 41% on average.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Inherent Trade-Offs between Diversity and Stability in Multi-Task BenchmarksGuanhua Zhang, Moritz HardtICML 2024 · 被引用 22 次
- ILDAE: Instance-Level Difficulty Analysis of Evaluation DataNeeraj Varshney, Swaroop Mishra, Chitta BaralACL 2022 · 被引用 21 次
- PMU Tracker: A Visualization Platform for Epicentric Event Propagation Analysis in the Power GridAnjana Arunkumar, Andrea Pinceti, Lalitha Sankar, Chris BryanIEEE VIS 2022 · 被引用 9 次
- ECBD: Evidence-Centered Benchmark Design for NLPYu Lu Liu, Su Lin Blodgett, Jackie C. K. Cheung, Vera Liao 等ACL 2024 · 被引用 3 次
- Metritocracy: Representative Metrics for Lite BenchmarksAriel D. Procaccia, Ben Schiffer, Serena Wang, Shirley ZhangNeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper6
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
- Adversarial Filters of Dataset BiasesRonan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers 等ICML 2020 · 被引用 242 次
- Selective Question Answering under Domain ShiftAmita Kamath, Robin Jia, Percy LiangACL 2020 · 被引用 121 次
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 被引用 51 次
相关 Paper
- Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?Pedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor 等ACL 2021
- Exploiting Leaderboards for Large-Scale Distribution of Malicious ModelsAnshuman Suri, Harsh Chaudhari, Yuefeng Peng, Ali Naseh 等S&P 2026
- When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model LeaderboardsNorah A. Alzahrani, Hisham Abdullah Alyahya, Yazeed Alnumay, Sultan Alrashed 等ACL 2024 · 被引用 13 次
- Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent EvaluationSayash Kapoor, Benedikt Stroebl, Peter Kirgis, Nitya Nadgir 等ICLR 2026 · 被引用 86 次
- Prompt-to-Leaderboard: Prompt-Adaptive LLM EvaluationsEvan Frick, Connor Chen, Joseph Tennyson, Tianle Li 等ICML 2025
