Embracing Unimodal Aleatoric Uncertainty for Robust Multimodal Fusion
Zixian Gao, Xun Jiang, Xing Xu, Fumin Shen, Yujie Li, Heng Tao Shen
摘要
As a fundamental problem in multimodal learning, multimodal fusion aims to compensate for the inherent limitations of a single modality. One challenge of multimodal fusion is that the unimodal data in their unique embedding space mostly contains potential noise, which leads to corrupted cross-modal interactions. However, in this paper, we show that the potential noise in unimodal data could be well quantified and further employed to enhance more stable unimodal embeddings via contrastive learning. Specifically, we propose a novel generic and robust multimodal fusion strategy, termed Embracing Aleatoric Uncertainty (EAU), which is simple and can be applied to kinds of modalities. It consists of two key steps: (1) the Stable Unimodal Feature Augmentation (SUFA) that learns a stable unimodal representation by incorporating the aleatoric uncertainty into self-supervised contrastive learning. (2) Robust Multimodal Feature Integration (RMFI) leveraging an information-theoretic strategy to learn a robust compact joint representation. We evaluate our proposed EAU method on five multimodal datasets, where the video, RGB image, text, audio, and depth image are involved. Extensive experiments demonstrate the EAU method is more noiseresistant than existing multimodal fusion strategies and establishes new state-of-the-art on several benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Diffusion-Inspired Truncated Sampler for Text-Video RetrievalJiamian Wang, Pichao Wang, Dongfang Liu, Qiang Guan 等NeurIPS 2024 · 被引用 16 次
- Hyper-Modality Enhancement for Multimodal Sentiment Analysis with Missing ModalitiesYan Zhuang, Minhao Liu, Wei Bai, Yanru Zhang 等NeurIPS 2025 · 被引用 10 次
- CLCR: Cross-Level Semantic Collaborative Representation for Multimodal LearningChunlei Meng, Guanhong Huang, Rong Fu, Runmin Jian 等CVPR 2026 · 被引用 9 次
- Multimodal Negative LearningBaoquan Gong, Xiyuan Gao, Pengfei Zhu, Qinghua Hu 等NeurIPS 2025 · 被引用 4 次
- Plug-and-play Feature Causality Decomposition for Multimodal Representation LearningYe Liu, Zihan Ji, Hongmin CaiNeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper21
- Attention Bottlenecks for Multimodal FusionArsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen 等NeurIPS 2021 · 被引用 884 次
- Probabilistic Face EmbeddingsYichun Shi, Anil K. JainICCV 2019 · 被引用 362 次
- Balanced Multimodal Learning via On-the-fly Gradient ModulationXiaokang Peng, Yake Wei, Andong Deng, Dong Wang 等CVPR 2022 · 被引用 264 次
- Modality Competition: What Makes Joint Training of Multi-modal Network Fail in Deep Learning? (Provably)Yu Huang, Junyang Lin, Chang Zhou, Hongxia Yang 等ICML 2022 · 被引用 168 次
- Provable Dynamic Fusion for Low-Quality Multimodal DataQingyang Zhang, Haitao Wu, Changqing Zhang, Qinghua Hu 等ICML 2023 · 被引用 143 次
相关 Paper
- Robust Multi-View Learning via Representation Fusion of Sample-Level Attention and Alignment of Simulated PerturbationJie Xu, Na Zhao, Gang Niu, Masashi Sugiyama 等ICCV 2025 · 被引用 6 次
- EquiAV: Leveraging Equivariance for Audio-Visual Contrastive LearningJongsuk Kim, Hyeongkeun Lee, Kyeongha Rho, Junmo Kim 等ICML 2024 · 被引用 15 次
- LAUA: Handling Missing Modalities and Unpaired Data in Multimodal Federated LearningYi Wei, Xiaokai Zhou, Shanshan Feng, Chuang Hu 等KDD 2026
- THE MORE, THE MERRIER: CONTRASTIVE FUSION FOR HIGHER-ORDER MULTIMODAL ALIGNMENTStefanos Koutoupis, Michaela Areti Zervou, Konstantinos Kontras, Maarten De Vos 等CVPR 2026 · 被引用 5 次
- AVFF: Audio-Visual Feature Fusion for Video Deepfake DetectionTrevine Oorloff, Surya Koppisetti, Nicolò Bonettini, Divyaraj Solanki 等CVPR 2024 · 被引用 51 次
