Prompt-to-Leaderboard: Prompt-Adaptive LLM Evaluations
Evan Frick, Connor Chen, Joseph Tennyson, Tianle Li, Wei-Lin Chiang, Anastasios Nikolas Angelopoulos, Ion Stoica
Abstract
Large language model (LLM) evaluations typically rely on aggregated metrics like accuracy or human preference, averaging across users and prompts. This averaging obscures userand prompt-specific variations in model performance. To address this, we propose Promptto-Leaderboard (P2L), a method that produces leaderboards specific to a prompt or set of prompts. The core idea is to train an LLM taking natural language prompts as input to output a vector of Bradley-Terry coefficients which are then used to predict the human preference vote. The resulting prompt-dependent leaderboards allow for unsupervised task-specific evaluation, optimal routing of queries to models, personalization, and automated evaluation of model strengths and weaknesses. Data from Chatbot Arena suggest that P2L better captures the nuanced landscape of language model performance than the averaged leaderboard. Furthermore, our findings suggest that P2L's ability to produce prompt-specific evaluations follows a power law scaling similar to that observed in LLMs themselves. In January 2025, the router we trained based on this methodology achieved the #1 spot on the Chatbot Arena leaderboard. Our code is available at this GitHub link: https://github.com/lmarena/p2l .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b24e9427-c3dd-422a-9be2-76dfe0f7a47eCited by top-tier papers4
- Nonparametric LLM Evaluation from Preference DataDennis Frauen, Athiya Deviyani, Mihaela van der Schaar, Stefan FeuerriegelICML 2026 · 5 citations
- Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative SignalsZihan Dong, Zhixian Zhang, Yang Zhou, Can Jin et al.ICML 2026 · 2 citations
- Among Us: Measuring and Mitigating Malicious Contributions in Model Collaboration SystemsZiyuan Yang, Wenxuan Ding, Shangbin Feng, Yulia TsvetkovACL 2026 · 1 citation
- Pluralistic LeaderboardsNika Haghtalab, Ariel Procaccia, Han Shao, Serena Wang et al.ICML 2026
Builds on8
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Hybrid LLM: Cost-Efficient and Quality-Aware Query RoutingDujian Ding, Ankur Mallick, Chi Wang, Robert Sim et al.ICLR 2024 · 282 citations
- Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise ComparisonsBanghua Zhu, Michael I. Jordan, Jiantao JiaoICML 2023 · 273 citations
- tinyBenchmarks: evaluating LLMs with fewer examplesFelipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun et al.ICML 2024 · 212 citations
Related papers
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- Efficient multi-prompt evaluation of LLMsFelipe Maia Polo, Ronald Xu, Lucas Weber, Mírian Silva et al.NeurIPS 2024 · 93 citations
- Copilot Arena: A Platform for Code LLM Evaluation in the WildWayne Chi, Valerie Chen, Anastasios Nikolas Angelopoulos, Wei-Lin Chiang et al.ICML 2025
- Exploring and Mitigating Adversarial Manipulation of Voting-Based LeaderboardsYangsibo Huang, Milad Nasr, Anastasios Nikolas Angelopoulos, Nicholas Carlini et al.ICML 2025
- From Crowdsourced Data to High-quality Benchmarks: Arena-Hard and Benchbuilder PipelineTianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap et al.ICML 2025
