MULTIBENCH++: A Unified and Comprehensive Multimodal Fusion Benchmarking Across Specialized Domains
Leyan Xue, Changqing Zhang, Kecheng Xue, Xiaohong Liu, Guangyu Wang, Zongbo Han
Abstract
Although multimodal fusion has made significant progress, its advancement is severely hindered by the lack of adequate evaluation benchmarks. Current fusion methods are typically evaluated on a small selection of public datasets, a limited scope that inadequately represents the complexity and diversity of real-world scenarios, potentially leading to biased evaluations. This issue presents a twofold challenge. On one hand, models may overfit to the biases of specific datasets, hindering their generalization to broader practical applications. On the other hand, the absence of a unified evaluation standard makes fair and objective comparisons between different fusion methods difficult. Consequently, a truly universal and high-performance fusion model has yet to emerge. To address these challenges, we have developed a large-scale, domain-adaptive benchmark for multimodal evaluation. This benchmark integrates over 30 datasets, encompassing 15 modalities and 20 predictive tasks across key application domains. To complement this, we have also developed an open-source, unified, and automated evaluation pipeline that includes standardized implementations of state-of-the-art models and diverse fusion paradigms. Leveraging this platform, we have conducted large-scale experiments, successfully establishing new performance baselines across multiple tasks. This work provides the academic community with a crucial platform for rigorous and reproducible assessment of multimodal models, aiming to propel the field of multimodal artificial intelligence to new heights.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 840b8bef-cbb3-4c58-9646-c2bca1ee31dbBuilds on10
- Attention Bottlenecks for Multimodal FusionArsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen et al.NeurIPS 2021 · 884 citations
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 701 citations
- CH-SIMS: A Chinese Multimodal Sentiment Analysis Dataset with Fine-grained Annotation of ModalityWenmeng Yu, Hua Xu, Fanyang Meng, Yilin Zhu et al.ACL 2020 · 376 citations
- Multimodal Token Fusion for Vision TransformersYikai Wang, Xinghao Chen, Lele Cao, Wenbing Huang et al.CVPR 2022 · 214 citations
- Provable Dynamic Fusion for Low-Quality Multimodal DataQingyang Zhang, Haitao Wu, Changqing Zhang, Qinghua Hu et al.ICML 2023 · 143 citations
Related papers
- Unison: Benchmarking Unified Multimodal Models via Synergistic Understanding and GenerationJinyu Liu, Xincheng Shuai, Henghui Ding, Yu-Gang JiangICML 2026 · 2 citations
- MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation ModelsWulin Xie, YiFan Zhang, Chaoyou Fu, Yang Shi et al.ICLR 2026 · 31 citations
- USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language ModelsBaolin Zheng, Guanlin Chen, Qingyang Teng, Hongqiong Zhong et al.ACL 2026 · 10 citations
- MixEval-X: Any-to-any Evaluations from Real-world Data MixtureJinjie Ni, Yifan Song, Deepanway Ghosal, Bo Li et al.ICLR 2025
- UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual DocumentsYifan Ji, Zhipeng Xu, Zhenghao Liu, Zulong Chen et al.ACL 2026 · 3 citations
