Embracing Unimodal Aleatoric Uncertainty for Robust Multimodal Fusion
Zixian Gao, Xun Jiang, Xing Xu, Fumin Shen, Yujie Li, Heng Tao Shen
Abstract
As a fundamental problem in multimodal learning, multimodal fusion aims to compensate for the inherent limitations of a single modality. One challenge of multimodal fusion is that the unimodal data in their unique embedding space mostly contains potential noise, which leads to corrupted cross-modal interactions. However, in this paper, we show that the potential noise in unimodal data could be well quantified and further employed to enhance more stable unimodal embeddings via contrastive learning. Specifically, we propose a novel generic and robust multimodal fusion strategy, termed Embracing Aleatoric Uncertainty (EAU), which is simple and can be applied to kinds of modalities. It consists of two key steps: (1) the Stable Unimodal Feature Augmentation (SUFA) that learns a stable unimodal representation by incorporating the aleatoric uncertainty into self-supervised contrastive learning. (2) Robust Multimodal Feature Integration (RMFI) leveraging an information-theoretic strategy to learn a robust compact joint representation. We evaluate our proposed EAU method on five multimodal datasets, where the video, RGB image, text, audio, and depth image are involved. Extensive experiments demonstrate the EAU method is more noiseresistant than existing multimodal fusion strategies and establishes new state-of-the-art on several benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c952e1a-72e5-4549-bce3-e9462986d30bCited by top-tier papers22
- Diffusion-Inspired Truncated Sampler for Text-Video RetrievalJiamian Wang, Pichao Wang, Dongfang Liu, Qiang Guan et al.NeurIPS 2024 · 16 citations
- Hyper-Modality Enhancement for Multimodal Sentiment Analysis with Missing ModalitiesYan Zhuang, Minhao Liu, Wei Bai, Yanru Zhang et al.NeurIPS 2025 · 10 citations
- CLCR: Cross-Level Semantic Collaborative Representation for Multimodal LearningChunlei Meng, Guanhong Huang, Rong Fu, Runmin Jian et al.CVPR 2026 · 9 citations
- Multimodal Negative LearningBaoquan Gong, Xiyuan Gao, Pengfei Zhu, Qinghua Hu et al.NeurIPS 2025 · 4 citations
- Plug-and-play Feature Causality Decomposition for Multimodal Representation LearningYe Liu, Zihan Ji, Hongmin CaiNeurIPS 2025 · 4 citations
Builds on21
- Attention Bottlenecks for Multimodal FusionArsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen et al.NeurIPS 2021 · 884 citations
- Probabilistic Face EmbeddingsYichun Shi, Anil K. JainICCV 2019 · 362 citations
- Balanced Multimodal Learning via On-the-fly Gradient ModulationXiaokang Peng, Yake Wei, Andong Deng, Dong Wang et al.CVPR 2022 · 264 citations
- Modality Competition: What Makes Joint Training of Multi-modal Network Fail in Deep Learning? (Provably)Yu Huang, Junyang Lin, Chang Zhou, Hongxia Yang et al.ICML 2022 · 168 citations
- Provable Dynamic Fusion for Low-Quality Multimodal DataQingyang Zhang, Haitao Wu, Changqing Zhang, Qinghua Hu et al.ICML 2023 · 143 citations
Related papers
- Robust Multi-View Learning via Representation Fusion of Sample-Level Attention and Alignment of Simulated PerturbationJie Xu, Na Zhao, Gang Niu, Masashi Sugiyama et al.ICCV 2025 · 6 citations
- EquiAV: Leveraging Equivariance for Audio-Visual Contrastive LearningJongsuk Kim, Hyeongkeun Lee, Kyeongha Rho, Junmo Kim et al.ICML 2024 · 15 citations
- LAUA: Handling Missing Modalities and Unpaired Data in Multimodal Federated LearningYi Wei, Xiaokai Zhou, Shanshan Feng, Chuang Hu et al.KDD 2026
- THE MORE, THE MERRIER: CONTRASTIVE FUSION FOR HIGHER-ORDER MULTIMODAL ALIGNMENTStefanos Koutoupis, Michaela Areti Zervou, Konstantinos Kontras, Maarten De Vos et al.CVPR 2026 · 5 citations
- AVFF: Audio-Visual Feature Fusion for Video Deepfake DetectionTrevine Oorloff, Surya Koppisetti, Nicolò Bonettini, Divyaraj Solanki et al.CVPR 2024 · 51 citations
