ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities
Adhiraj Ghosh, Sebastian Dziadzio, Ameya Prabhu, Vishaal Udandarao, Samuel Albanie, Matthias Bethge
摘要
Traditional fixed test datasets fall short in evaluating the open-ended capabilities of foundation models. To address this, we propose ONEBench (OpeN-Ended Benchmarking), a new paradigm that consolidates individual evaluation datasets into a unified, ever-expanding sample pool. ONEBench enables custom benchmarks for specific capabilities while reusing and aggregating samples, mitigating overfitting and dataset bias for broader capability assessment. It reframes model evaluation as selecting and aggregating samplelevel tests. Transitioning from task-specific benchmarks to ONEBench introduces two challenges: heterogeneity (aggregating diverse metrics) and incompleteness (comparing models tested on different data subsets). To address these, we propose an aggregation algorithm that ensures identifiability-asymptotically recovering ground-truth scores-and rapid convergence, enabling accurate model comparisons with relatively little data. On homogenous datasets, our algorithm produces rankings that highly correlate with average scores. Moreover, it remains robust to over 95% missing measurements, reducing evaluation costs by up to 20 times. We introduce ONEBench-LLM for language models and ONEBench-LMM for visionlanguage models, enabling targeted model testing across diverse capabilities. * Equal contribution, random order, • core contributors 1 From a talk by Alexei Efros at ICML 2020
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- What Moves the Eyes: Doubling Mechanistic Model Performance Using Deep Networks to Discover and Test Cognitive HypothesesFederico D'Agostino, Lisa Schwetlick, Matthias Bethge, Matthias KümmererNeurIPS 2025 · 被引用 4 次
- Concept-Aware Batch Sampling Improves Language-Image PretrainingAdhiraj Ghosh, Vishaal Udandarao, Thao Nguyen, Matteo Farina 等CVPR 2026 · 被引用 1 次
- Great Models Think Alike and this Undermines AI OversightShashwat Goel, Joschka Strüber, Ilze Amanda Auzina, Karuna K. Chandra 等ICML 2025
- Probing the Knowledge Boundary: An Interactive Agentic Framework for Deep Knowledge ExtractionYuheng Yang, Siqi Zhu, Tao Feng, Ge Liu 等ICML 2026
- WikiBigEdit: Understanding the Limits of Lifelong Knowledge Editing in LLMsLukas Thede, Karsten Roth, Matthias Bethge, Zeynep Akata 等ICML 2025
它引用的顶会 Paper28
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 等NeurIPS 2022 · 被引用 2,727 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang 等ICML 2024 · 被引用 1,191 次
相关 Paper
- MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World TasksJiacheng Chen, Tianhao Liang, Sherman Siu, Zhengqing Wang 等ICLR 2025
- UniVBench: Towards Unified Evaluation for Video Foundation ModelsJianhui Wei, Xiaotian Zhang, Yichen Li, Yuan Wang 等CVPR 2026 · 被引用 12 次
- How Benchmark Prediction from Fewer Data Misses the MarkGuanhua Zhang, Florian E. Dorner, Moritz HardtNeurIPS 2025 · 被引用 26 次
- VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning TasksShiduo Zhang, Zhe Xu, Peiju Liu, Xiaopeng Yu 等ICCV 2025 · 被引用 12 次
- On the Evaluation of Capability Estimation Methods for Large Language ModelsQiang Hu, Jin Wen, Yao Zhang, Maxime Cordy 等AAAI 2026
