Dynamic Multimodal Evaluation via Knowledge-Enhanced Benchmark Evolution
Junzhe Zhang, Huixuan Zhang, Xiaojun Wan
摘要
The rapid development of multimodal large language models (MLLMs) has created an urgent demand for more reliable and robust evaluation protocols, however, existing static benchmarks are prone to data contamination and performance saturation, which can result in inflated or misleading evaluation results. To address these limitations, we first introduce a graph formulation to represent both static and dynamic visual question answering (VQA) samples. Building upon this formulation, we propose Knowledge-Enhanced Benchmark Evolution (KBE), a dynamic multimodal evaluation framework that first analyzes the original static benchmark, then expands it by integrating multimodal knowledge, transforming the static benchmark into a controllable, dynamic evolving version. Crucially, KBE can both reconstruct questions by Re-selecting visual information in the original image and expand existing questions with external textual knowledge. By explicitly controlling the degree of question exploration, KBE enables difficulty-controllable evaluation across a wide range of model capabilities. Extensive experimental results demonstrate that KBE effectively mitigates data contamination and benchmark saturation, while providing a more comprehensive and flexible assessment of MLLM performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Are We on the Right Way for Evaluating Large Vision-Language Models?Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang 等NeurIPS 2024 · 被引用 1,029 次
- DyVal: Dynamic Evaluation of Large Language Models for Reasoning TasksKaijie Zhu, Jiaao Chen, Jindong Wang, Neil Zhenqiang Gong 等ICLR 2024 · 被引用 92 次
- NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity ClassesLizhou Fan, Wenyue Hua, Lingyao Li, Haoyang Ling 等ACL 2024 · 被引用 8 次
- Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving TestingHan Jiang, Xiaoyuan Yi, Zhihua Wei, Ziang Xiao 等ICML 2025
- Dynamic Multimodal Evaluation with Flexible Complexity by Vision-Language BootstrappingYue Yang, Shuibo Zhang, Kaipeng Zhang, Yi Bin 等ICLR 2025
相关 Paper
- MMBench-Live: A Continuously Evolving Benchmark for Multimodal ModelsYuanzhi Liu, Shousheng Zhao, Bo Zhou, Kongming Liang 等ICML 2026
- SDEval: Safety Dynamic Evaluation for Multimodal Large Language ModelsHanqing Wang, Yuan Tian, Mingyu Liu, Zhenhao Zhang 等AAAI 2026 · 被引用 2 次
- Benchmarking Deflection and Hallucination in Large Vision-Language ModelsNicholas Moratelli, Christopher Davis, Leonardo F. R. Ribeiro, Bill Byrne 等ACL 2026 · 被引用 1 次
- Hybrid-DMKG: A Hybrid Reasoning Framework over Dynamic Multimodal Knowledge Graphs for Multimodal Multihop QA with Knowledge EditingLi Yuan, Qingfei Huang, Bingshan Zhu, Yi Cai 等AAAI 2026
- VKG-QA: Visual Knowledge Graph-based Question Answer for Large Multimodal ModelsYuntao Du, Yiming Wang, Renshuo Yuan, Jincheng Yue 等CVPR 2026
