Improving Your Model Ranking on Chatbot Arena by Vote Rigging
Rui Min, Tianyu Pang, Chao Du, Qian Liu, Minhao Cheng, Min Lin
摘要
Chatbot Arena is a popular platform for evaluating LLMs by pairwise battles, where users vote for their preferred response from two randomly sampled anonymous models. While Chatbot Arena is widely regarded as a reliable LLM ranking leaderboard, we show that crowdsourced voting can be rigged to improve (or decrease) the ranking of a target model m t . We first introduce a straightforward target-only rigging strategy that focuses on new battles involving m t , identifying it via watermarking or a binary classifier, and exclusively voting for m t wins. However, this strategy is practically inefficient because there are over 190 models on Chatbot Arena and on average only about 1% of new battles will involve m t . To overcome this, we propose omnipresent rigging strategies, exploiting the Elo rating mechanism of Chatbot Arena that any new vote on a battle can influence the ranking of the target model m t , even if m t is not directly involved in the battle. We conduct experiments on around 1.7 million historical votes from the Chatbot Arena Notebook, showing that omnipresent rigging strategies can improve model rankings by rigging only hundreds of new votes. While we have evaluated several defense mechanisms, our findings highlight the importance of continued efforts to prevent vote rigging. Code is publicly available to reproduce all experiments. * Equal contribution. The project was done during Rui Min's internship at Sea AI Lab.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Dropping Just a Handful of Preferences Can Change Top Large Language Model RankingsJenny Y. Huang, Yunyi Shen, Dennis Wei, Tamara BroderickICLR 2026 · 被引用 8 次
- Strategic Candidacy in Generative AI ArenasChris Hays, Rachel Li, Bailey Flanigan, Manish RaghavanICML 2026 · 被引用 3 次
- Exploiting Leaderboards for Large-Scale Distribution of Malicious ModelsAnshuman Suri, Harsh Chaudhari, Yuefeng Peng, Ali Naseh 等S&P 2026
- Fewer Battles, More Gain: An Information-Efficient Framework for Arena-based LLM EvaluationZirui Liu, Xianquan Wang, Yan Zhuang, Jiatong Li 等ICLR 2026
- When Anonymity Breaks: Identifying Models Behind Text-to-Image LeaderboardsAli Naseh, Anshuman Suri, Yuefeng Peng, Harsh Chaudhari 等CVPR 2026
它引用的顶会 Paper16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz 等ICML 2023 · 被引用 854 次
相关 Paper
- Exploring and Mitigating Adversarial Manipulation of Voting-Based LeaderboardsYangsibo Huang, Milad Nasr, Anastasios Nikolas Angelopoulos, Nicholas Carlini 等ICML 2025
- Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee DiscussionsRuochen Zhao, Wenxuan Zhang, Yew Ken Chia, Weiwen Xu 等ACL 2025 · 被引用 34 次
- Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language ModelsYanbin Yin, Kun Zhou, Zhen Wang, Xiangdong Zhang 等ACL 2026 · 被引用 2 次
- Cheating Automatic LLM Benchmarks: Null Models Achieve High Win RatesXiaosen Zheng, Tianyu Pang, Chao Du, Qian Liu 等ICLR 2025
- WizardArena: Post-training Large Language Models via Simulated Offline Chatbot ArenaHaipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao 等NeurIPS 2024 · 被引用 7 次
