Unbiased Missing-Modality Multimodal Learning
Ruiting Dai, Chenxi Li, Yandong Yan, Lisi Mo, Ke Qin, Tao He
Abstract
Recovering missing modalities in multimodal learning has recently been approached using diffusion models to synthesize absent data conditioned on available modalities. However, existing methods often suffer from modality generation bias: while certain modalities are generated with high fidelity, others-such as video-remain challenging due to intrinsic modality gaps, leading to imbalanced training. To address this issue, we propose MD 2 N (Multi-stage Duplex Diffusion Network), a novel framework for unbiased missing-modality recovery. MD 2 N introduces a modality transfer module within a duplex diffusion architecture, enabling bidirectional generation between available and missing modalities through three stages: (1) global structure generation, (2) modality transfer, and (3) local crossmodal refinement. By training with duplex diffusion, both available and missing modalities generate each other in an intersecting manner, effectively achieving a balanced generation state. Extensive experiments demonstrate that MD 2 N significantly outperforms existing state-of-the-art methods, achieving up to 4% improvement over IMDer on the CMU-MOSEI dataset. Project page: here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 51b49040-d5e8-475b-85be-8027474187feCited by top-tier papers5
- PriorDrive: Enhancing Online HD Mapping with Unified Vector PriorsShuang Zeng, Xinyuan Chang, Xinran Liu, Yujian Yuan et al.AAAI 2026 · 12 citations
- TiCAL: Typicality-Based Consistency-Aware Learning for Multimodal Emotion RecognitionWen Yin, Siyu Zhan, Cencen Liu, Xin Hu et al.AAAI 2026 · 4 citations
- SepPrune: Structured Pruning for Efficient Deep Speech SeparationYuqi Li, Kai Li, Xin Yin, Zhifei Yang et al.AAAI 2026 · 4 citations
- AnyMod-LLVE: Low-Light Video Enhancement with Modality-Agnostic InferenceHangfeng Liang, Yutao Hu, Yanhan Hu, Xiaohan Wu et al.ICML 2026
- SPR: A Structured Prompt Refinement Network for Modality MissingHao Chen, Diwei Su, Zhuo Wang, Zuwang He et al.ICML 2026
Builds on18
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential EquationsChenlin Meng, Yutong He, Yang Song, Jiaming Song et al.ICLR 2022 · 2,128 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- SMIL: Multimodal Learning with Severely Missing ModalityMengmeng Ma, Jian Ren, Long Zhao, Sergey Tulyakov et al.AAAI 2021 · 393 citations
- A Diffusion-Based Framework for Multi-Class Anomaly DetectionHaoyang He, Jiangning Zhang, Hongxu Chen, Xuhai Chen et al.AAAI 2024 · 231 citations
- Incomplete Multimodality-Diffused Emotion RecognitionYuanzhi Wang, Yong Li, Zhen CuiNeurIPS 2023 · 155 citations
Related papers
- Rethinking Diffusion Bridge Model with Dual Alignments for Medical Image SynthesisJinbao Wei, Yuhang Chen, Zhijie Wang, Gang Yang et al.ACM MM 2025 · 3 citations
- Distribution-Consistent Modal Recovering for Incomplete Multimodal LearningYuanzhi Wang, Zhen Cui, Yong LiICCV 2023 · 101 citations
- MM-Align: Learning Optimal Transport-based Alignment Dynamics for Fast and Accurate Inference on Missing Modality SequencesWei Han, Hui Chen, Min-Yen Kan, Soujanya PoriaEMNLP 2022 · 13 citations
- Incomplete Cross-modal Retrieval with Dual-Aligned Variational AutoencodersMengmeng Jing, Jingjing Li, Lei Zhu, Ke Lu et al.ACM MM 2020 · 63 citations
- MaskMentor: Unlocking the Potential of Masked Self-Teaching for Missing Modality RGB-D Semantic SegmentationZhida Zhao, Jia Li, Lijun Wang, Yifan Wang et al.ACM MM 2024 · 1 citation
