Dynamic Multimodal Evaluation via Knowledge-Enhanced Benchmark Evolution
Junzhe Zhang, Huixuan Zhang, Xiaojun Wan
Abstract
The rapid development of multimodal large language models (MLLMs) has created an urgent demand for more reliable and robust evaluation protocols, however, existing static benchmarks are prone to data contamination and performance saturation, which can result in inflated or misleading evaluation results. To address these limitations, we first introduce a graph formulation to represent both static and dynamic visual question answering (VQA) samples. Building upon this formulation, we propose Knowledge-Enhanced Benchmark Evolution (KBE), a dynamic multimodal evaluation framework that first analyzes the original static benchmark, then expands it by integrating multimodal knowledge, transforming the static benchmark into a controllable, dynamic evolving version. Crucially, KBE can both reconstruct questions by Re-selecting visual information in the original image and expand existing questions with external textual knowledge. By explicitly controlling the degree of question exploration, KBE enables difficulty-controllable evaluation across a wide range of model capabilities. Extensive experimental results demonstrate that KBE effectively mitigates data contamination and benchmark saturation, while providing a more comprehensive and flexible assessment of MLLM performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 29f3b7cd-6271-41fe-80a5-0484c451f014Builds on5
- Are We on the Right Way for Evaluating Large Vision-Language Models?Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang et al.NeurIPS 2024 · 1,029 citations
- DyVal: Dynamic Evaluation of Large Language Models for Reasoning TasksKaijie Zhu, Jiaao Chen, Jindong Wang, Neil Zhenqiang Gong et al.ICLR 2024 · 92 citations
- NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity ClassesLizhou Fan, Wenyue Hua, Lingyao Li, Haoyang Ling et al.ACL 2024 · 8 citations
- Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving TestingHan Jiang, Xiaoyuan Yi, Zhihua Wei, Ziang Xiao et al.ICML 2025
- Dynamic Multimodal Evaluation with Flexible Complexity by Vision-Language BootstrappingYue Yang, Shuibo Zhang, Kaipeng Zhang, Yi Bin et al.ICLR 2025
Related papers
- MMBench-Live: A Continuously Evolving Benchmark for Multimodal ModelsYuanzhi Liu, Shousheng Zhao, Bo Zhou, Kongming Liang et al.ICML 2026
- SDEval: Safety Dynamic Evaluation for Multimodal Large Language ModelsHanqing Wang, Yuan Tian, Mingyu Liu, Zhenhao Zhang et al.AAAI 2026 · 2 citations
- Benchmarking Deflection and Hallucination in Large Vision-Language ModelsNicholas Moratelli, Christopher Davis, Leonardo F. R. Ribeiro, Bill Byrne et al.ACL 2026 · 1 citation
- Hybrid-DMKG: A Hybrid Reasoning Framework over Dynamic Multimodal Knowledge Graphs for Multimodal Multihop QA with Knowledge EditingLi Yuan, Qingfei Huang, Bingshan Zhu, Yi Cai et al.AAAI 2026
- VKG-QA: Visual Knowledge Graph-based Question Answer for Large Multimodal ModelsYuntao Du, Yiming Wang, Renshuo Yuan, Jincheng Yue et al.CVPR 2026
