Metritocracy: Representative Metrics for Lite Benchmarks
Ariel D. Procaccia, Ben Schiffer, Serena Wang, Shirley Zhang
Abstract
A common problem in LLM evaluation is how to choose a subset of metrics from a full suite of possible metrics. Subset selection is usually done for efficiency or interpretability reasons, and the goal is often to select a representative'' subset of metrics. However, representative'' is rarely clearly defined. In this work, we use ideas from social choice theory to formalize two notions of representation for the selection of a subset of evaluation metrics. We first introduce positional representation, which guarantees every alternative is sufficiently represented at every position cutoff. We then introduce positional proportionality, which guarantees no alternative is proportionally over- or under-represented by more than a small error at any position. We prove upper and lower bounds on the smallest number of metrics needed to guarantee either of these properties in the worst case. We also study a generalized form of each property that allows for additional input on groups of metrics that must be represented. Finally, we tie theory to practice through real-world case studies on both LLM evaluation and hospital quality evaluation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5079ebbe-fd8c-4a90-ac96-65c5cfe02ca5Builds on10
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- tinyBenchmarks: evaluating LLMs with fewer examplesFelipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun et al.ICML 2024 · 212 citations
- Proportional Participatory Budgeting with Additive UtilitiesDominik Peters, Grzegorz Pierczynski, Piotr SkowronNeurIPS 2021 · 168 citations
- InfoLM: A New Metric to Evaluate Summarization & Data2Text GenerationPierre Jean A. Colombo, Chloé Clavel, Pablo PiantanidaAAAI 2022 · 52 citations
- How Robust are Model Rankings : A Leaderboard Customization Approach for Equitable EvaluationSwaroop Mishra, Anjana ArunkumarAAAI 2021 · 27 citations
Related papers
- Is Sortition Both Representative and Fair?Soroush Ebadian, Gregory Kehne, Evi Micha, Ariel D. Procaccia et al.NeurIPS 2022 · 24 citations
- Comparing Election Methods Where Each Voter Ranks Only Few CandidatesMatthias Bentert, Piotr SkowronAAAI 2020 · 21 citations
- Can a Few Decide for Many? The Metric Distortion of SortitionIoannis Caragiannis, Evi Micha, Jannik PetersICML 2024 · 11 citations
- Worst-Case Voting When the Stakes Are HighAnson Kahng, Gregory KehneAAAI 2022 · 3 citations
- Facility Location for Fair and Equitable Query ResultsSara Cohen, Helen SternbachICDE 2025
