Redundancy Principles for MLLMs Benchmarks
Zicheng Zhang, Xiangyu Zhao, Xinyu Fang, Chunyi Li, Xiaohong Liu, Xiongkuo Min, Haodong Duan, Kai Chen, Guangtao Zhai
Abstract
With the rapid iteration of Multi-modality Large Language Models (MLLMs) and the evolving demands of the field, the number of benchmarks produced annually has surged into the hundreds. The rapid growth has inevitably led to significant redundancy among benchmarks. Therefore, it is crucial to take a step back and critically assess the current state of redundancy and propose targeted principles for constructing effective MLLM benchmarks. In this paper, we focus on redundancy from three key perspectives: 1) Redundancy of benchmark capability dimensions, 2) Redundancy in the number of test questions, and 3) Cross-benchmark redundancy within specific domains. Through the comprehensive analysis over hundreds of MLLMs' performance across more than 20 benchmarks, we aim to quantitatively measure the level of redundancy lies in existing MLLM evaluations, provide valuable insights to guide the future development of MLLM benchmarks, and offer strategies to refine and address redundancy issues effectively. The code is available at https://github.com/zzc-1998/Benchmark-Redundancy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a04d642d-ce57-4510-9359-abaa032ca57bCited by top-tier papers8
- Creation-Mmbench: Assessing Context-Aware Creative Intelligence in MllmsXinyu Fang, Zhijian Chen, Kai Lan, Lixin Ma et al.ICCV 2025 · 23 citations
- Rethinking LLM Evaluation: Can We Evaluate LLMs with 200× Less Data?Shaobo Wang, Cong Wang, Wenjie Fu, Yue Min et al.ICLR 2026 · 2 citations
- OSCBench: Benchmarking Object State Change in Text-to-Video GenerationXianjing Han, Bin Zhu, Shiqi Hu, Franklin Mingzhe Li et al.ACL 2026 · 2 citations
- Information Density Principle for MLLM BenchmarksChunyi Li, Xiaozhe Li, Zicheng Zhang, Yuan Tian et al.ICCV 2025 · 1 citation
- MetaEval: Measuring the Discrimination of Benchmarks for Efficient LLM EvaluationZhuo Wang, Wen Wu, Guoqing Wang, Guangze Ye et al.AAAI 2026 · 1 citation
Builds on10
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu et al.ICLR 2024 · 1,472 citations
- GPT4Tools: Teaching Large Language Model to Use Tools via Self-instructionRui Yang, Lin Song, Yanwei Li, Sijie Zhao et al.NeurIPS 2023 · 340 citations
- Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level VisionHaoning Wu, Zicheng Zhang, Erli Zhang, Chaofeng Chen et al.ICLR 2024 · 258 citations
Related papers
- Evaluating Cross-Modal Reasoning Ability and Problem Characteristics with Multimodal Item Response TheoryShunki Uebayashi, Kento Masui, Kyohei Atarashi, Han Bao et al.ICLR 2026 · 1 citation
- tinyBenchmarks: evaluating LLMs with fewer examplesFelipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun et al.ICML 2024 · 212 citations
- On Path to Multimodal Generalist: General-Level and General-BenchHao Fei, Yuan Zhou, Juncheng Li, Xiangtai Li et al.ICML 2025
- LLM-Powered Benchmark Factory: Reliable, Generic, and EfficientPeiwen Yuan, Shaoxiong Feng, Yiwei Li, Xinglin Wang et al.ACL 2026 · 8 citations
- HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized BenchmarksTing Zhou, Daoyuan Chen, Qirui Jiao, Bolin Ding et al.CVPR 2026
