ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities
Adhiraj Ghosh, Sebastian Dziadzio, Ameya Prabhu, Vishaal Udandarao, Samuel Albanie, Matthias Bethge
Abstract
Traditional fixed test datasets fall short in evaluating the open-ended capabilities of foundation models. To address this, we propose ONEBench (OpeN-Ended Benchmarking), a new paradigm that consolidates individual evaluation datasets into a unified, ever-expanding sample pool. ONEBench enables custom benchmarks for specific capabilities while reusing and aggregating samples, mitigating overfitting and dataset bias for broader capability assessment. It reframes model evaluation as selecting and aggregating samplelevel tests. Transitioning from task-specific benchmarks to ONEBench introduces two challenges: heterogeneity (aggregating diverse metrics) and incompleteness (comparing models tested on different data subsets). To address these, we propose an aggregation algorithm that ensures identifiability-asymptotically recovering ground-truth scores-and rapid convergence, enabling accurate model comparisons with relatively little data. On homogenous datasets, our algorithm produces rankings that highly correlate with average scores. Moreover, it remains robust to over 95% missing measurements, reducing evaluation costs by up to 20 times. We introduce ONEBench-LLM for language models and ONEBench-LMM for visionlanguage models, enabling targeted model testing across diverse capabilities. * Equal contribution, random order, • core contributors 1 From a talk by Alexei Efros at ICML 2020
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d71198b4-1436-4915-83ec-ad5d30491cdeCited by top-tier papers5
- What Moves the Eyes: Doubling Mechanistic Model Performance Using Deep Networks to Discover and Test Cognitive HypothesesFederico D'Agostino, Lisa Schwetlick, Matthias Bethge, Matthias KümmererNeurIPS 2025 · 4 citations
- Concept-Aware Batch Sampling Improves Language-Image PretrainingAdhiraj Ghosh, Vishaal Udandarao, Thao Nguyen, Matteo Farina et al.CVPR 2026 · 1 citation
- Great Models Think Alike and this Undermines AI OversightShashwat Goel, Joschka Strüber, Ilze Amanda Auzina, Karuna K. Chandra et al.ICML 2025
- Probing the Knowledge Boundary: An Interactive Agentic Framework for Deep Knowledge ExtractionYuheng Yang, Siqi Zhu, Tao Feng, Ge Liu et al.ICML 2026
- WikiBigEdit: Understanding the Limits of Lifelong Knowledge Editing in LLMsLukas Thede, Karsten Roth, Matthias Bethge, Zeynep Akata et al.ICML 2025
Builds on28
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang et al.ICML 2024 · 1,191 citations
Related papers
- MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World TasksJiacheng Chen, Tianhao Liang, Sherman Siu, Zhengqing Wang et al.ICLR 2025
- UniVBench: Towards Unified Evaluation for Video Foundation ModelsJianhui Wei, Xiaotian Zhang, Yichen Li, Yuan Wang et al.CVPR 2026 · 12 citations
- How Benchmark Prediction from Fewer Data Misses the MarkGuanhua Zhang, Florian E. Dorner, Moritz HardtNeurIPS 2025 · 26 citations
- VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning TasksShiduo Zhang, Zhe Xu, Peiju Liu, Xiaopeng Yu et al.ICCV 2025 · 12 citations
- On the Evaluation of Capability Estimation Methods for Large Language ModelsQiang Hu, Jin Wen, Yao Zhang, Maxime Cordy et al.AAAI 2026
