K-Sort Arena: Efficient and Reliable Benchmarking for Generative Models via K-wise Human Preferences
Zhikai Li, Xuewen Liu, Dongrong Joe Fu, Jianquan Li, Qingyi Gu, Kurt Keutzer, Zhen Dong
Abstract
The rapid advancement of visual generative models necessitates efficient and reliable evaluation methods. Arena platform, which gathers user votes on model comparisons, can rank models with human preferences. However, traditional Arena methods, while established, require an excessive number of comparisons for ranking to converge and are vulnerable to preference noise in voting, suggesting the need for better approaches tailored to contemporary evaluation challenges. In this paper, we introduce K-Sort Arena, an efficient and reliable platform based on a key insight: images and videos possess higher perceptual intuitiveness than texts, enabling rapid evaluation of multiple samples simultaneously. Consequently, K-Sort Arena employs Kwise comparisons, allowing K models to engage in free-forall competitions, which yield much richer information than pairwise comparisons. To enhance the robustness of the system, we leverage probabilistic modeling and Bayesian updating techniques. We propose an exploration-exploitationbased matchmaking strategy to facilitate more informative comparisons. In our experiments, K-Sort Arena exhibits 16.3× faster convergence compared to the widely used ELO algorithm. To further validate the superiority and obtain a comprehensive leaderboard, we collect human feedback via crowdsourced evaluations of numerous cutting-edge textto-image and text-to-video models. Thanks to its high efficiency, K-Sort Arena can continuously incorporate emerging models and update the leaderboard with minimal votes. Our project has undergone several months of internal testing and is now available at K-Sort Arena.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ff72cb7-f0e7-4ca8-8fbf-aa19a0f9c7a5Cited by top-tier papers8
- WorldLens: Full-Spectrum Evaluations of Driving World Models in Real WorldAo Liang, Lingdong Kong, Tianyi Yan, Hongsi Liu et al.CVPR 2026 · 28 citations
- K-Sort Eval: Efficient Preference Evaluation for Visual Generation via Corrected VLM-as-a-JudgeZhikai Li, Jiatong Li, Xuewen Liu, Wangbo Zhao et al.ICLR 2026 · 3 citations
- PTQ4ARVG: Post-Training Quantization for AutoRegressive Visual Generation ModelsXuewen Liu, Zhikai Li, Jing Zhang, Mengjuan Chen et al.ICLR 2026 · 3 citations
- CacheQuant: Comprehensively Accelerated Diffusion ModelsXuewen Liu, Zhikai Li, Qingyi GuCVPR 2025
- OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text GenerationPengfei Zhou, Xiaopeng Peng, Jiajun Song, Chuanhao Li et al.CVPR 2025
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D GenerationTong Wu, Guandao Yang, Zhibing Li, Kai Zhang et al.CVPR 2024 · 42 citations
- Fewer Battles, More Gain: An Information-Efficient Framework for Arena-based LLM EvaluationZirui Liu, Xianquan Wang, Yan Zhuang, Jiatong Li et al.ICLR 2026
- am-ELO: A Stable Framework for Arena-based LLM EvaluationZirui Liu, Jiatong Li, Yan Zhuang, Qi Liu et al.ICML 2025
- GeoArena: Evaluating Open-World Geographic Reasoning in Large Vision-Language ModelsPengyue Jia, Yingyi Zhang, Xiangyu Zhao, Sharon LiACL 2026 · 3 citations
