Metritocracy: Representative Metrics for Lite Benchmarks
Ariel D. Procaccia, Ben Schiffer, Serena Wang, Shirley Zhang
摘要
A common problem in LLM evaluation is how to choose a subset of metrics from a full suite of possible metrics. Subset selection is usually done for efficiency or interpretability reasons, and the goal is often to select a representative'' subset of metrics. However, representative'' is rarely clearly defined. In this work, we use ideas from social choice theory to formalize two notions of representation for the selection of a subset of evaluation metrics. We first introduce positional representation, which guarantees every alternative is sufficiently represented at every position cutoff. We then introduce positional proportionality, which guarantees no alternative is proportionally over- or under-represented by more than a small error at any position. We prove upper and lower bounds on the smallest number of metrics needed to guarantee either of these properties in the worst case. We also study a generalized form of each property that allows for additional input on groups of metrics that must be represented. Finally, we tie theory to practice through real-world case studies on both LLM evaluation and hospital quality evaluation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- tinyBenchmarks: evaluating LLMs with fewer examplesFelipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun 等ICML 2024 · 被引用 212 次
- Proportional Participatory Budgeting with Additive UtilitiesDominik Peters, Grzegorz Pierczynski, Piotr SkowronNeurIPS 2021 · 被引用 168 次
- InfoLM: A New Metric to Evaluate Summarization & Data2Text GenerationPierre Jean A. Colombo, Chloé Clavel, Pablo PiantanidaAAAI 2022 · 被引用 52 次
- How Robust are Model Rankings : A Leaderboard Customization Approach for Equitable EvaluationSwaroop Mishra, Anjana ArunkumarAAAI 2021 · 被引用 27 次
相关 Paper
- Is Sortition Both Representative and Fair?Soroush Ebadian, Gregory Kehne, Evi Micha, Ariel D. Procaccia 等NeurIPS 2022 · 被引用 24 次
- Comparing Election Methods Where Each Voter Ranks Only Few CandidatesMatthias Bentert, Piotr SkowronAAAI 2020 · 被引用 21 次
- Can a Few Decide for Many? The Metric Distortion of SortitionIoannis Caragiannis, Evi Micha, Jannik PetersICML 2024 · 被引用 11 次
- Worst-Case Voting When the Stakes Are HighAnson Kahng, Gregory KehneAAAI 2022 · 被引用 3 次
- Facility Location for Fair and Equitable Query ResultsSara Cohen, Helen SternbachICDE 2025
