Provable Dynamic Fusion for Low-Quality Multimodal Data
Qingyang Zhang, Haitao Wu, Changqing Zhang, Qinghua Hu, Huazhu Fu, Joey Tianyi Zhou, Xi Peng
Abstract
The inherent challenge of multimodal fusion is to precisely capture the cross-modal correlation and flexibly conduct cross-modal interaction. To fully release the value of each modality and mitigate the influence of low-quality multimodal data, dynamic multimodal fusion emerges as a promising learning paradigm. Despite its widespread use, theoretical justifications in this field are still notably lacking. Can we design a provably robust multimodal fusion method? This paper provides theoretical understandings to answer this question under a most popular multimodal fusion framework from the generalization perspective. We proceed to reveal that several uncertainty estimation solutions are naturally available to achieve robust multimodal fusion. Then a novel multimodal fusion framework termed Quality-aware Multimodal Fusion (QMF) is proposed, which can improve the performance in terms of classification accuracy and model robustness. Extensive experimental results on multiple benchmarks can support our findings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 83c8228d-5117-48e5-a6e4-b2776ee2f40fCited by top-tier papers51
- Reliable Conflictive Multi-View LearningCai Xu, Jiajun Si, Ziyu Guan, Wei Zhao et al.AAAI 2024 · 121 citations
- Facilitating Multimodal Classification via Dynamically Learning Modality GapYang Yang, Fengqiang Wan, Qing-Yuan Jiang, Yi XuNeurIPS 2024 · 65 citations
- Task-Customized Mixture of Adapters for General Image FusionPengfei Zhu, Yang Sun, Bing Cao, Qinghua HuCVPR 2024 · 58 citations
- Classifier-guided Gradient Modulation for Enhanced Multimodal LearningZirun Guo, Tao Jin, Jingyuan Chen, Zhou ZhaoNeurIPS 2024 · 56 citations
- ReconBoost: Boosting Can Achieve Modality ReconcilementCong Hua, Qianqian Xu, Shilong Bao, Zhiyong Yang et al.ICML 2024 · 52 citations
Builds on26
- Energy-based Out-of-distribution DetectionWeitang Liu, Xiaoyun Wang, John D. Owens, Yixuan LiNeurIPS 2020 · 2,213 citations
- What Makes Multi-Modal Learning Better than Single (Provably)Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen et al.NeurIPS 2021 · 404 citations
- Balanced Multimodal Learning via On-the-fly Gradient ModulationXiaokang Peng, Yake Wei, Andong Deng, Dong Wang et al.CVPR 2022 · 264 citations
- Posterior Network: Uncertainty Estimation without OOD Samples via Density-Based Pseudo-CountsBertrand Charpentier, Daniel Zügner, Stephan GünnemannNeurIPS 2020 · 263 citations
- Training independent subnetworks for robust predictionMarton Havasi, Rodolphe Jenatton, Stanislav Fort, Jeremiah Zhe Liu et al.ICLR 2021 · 235 citations
Related papers
- Predictive Dynamic FusionBing Cao, Yinan Xia, Yi Ding, Changqing Zhang et al.ICML 2024 · 31 citations
- A Theoretical Proof of Dynamic Multimodal Fusion Exacerbates Modality GreedyXiaorui Ding, Huan Ma, Changqing ZhangACM MM 2025
- Proxy-Driven Robust Multimodal Sentiment Analysis with Incomplete DataAoqiang Zhu, Min Hu, Xiaohua Wang, Jiaoyun Yang et al.ACL 2025 · 8 citations
- CMoB: Modality Valuation via Causal Effect for Balanced Multimodal LearningJun Wang, Fuyuan Cao, Zhixin Xue, Xingwang Zhao et al.NeurIPS 2025 · 4 citations
- Embracing Unimodal Aleatoric Uncertainty for Robust Multimodal FusionZixian Gao, Xun Jiang, Xing Xu, Fumin Shen et al.CVPR 2024
