AOEPT: Breaking the Implicit Modality-Reduction Bottleneck in Modality-Missing Prompt Tuning
Jian Lang, Hong, Ting Zhong, Fan Zhou
Abstract
Deploying multimodal systems in real-world environments often entails handling modality-missing scenarios, where one or more modalities are unavailable. While recent studies address this challenge for the general Multimodal Transformer (MT) architecture via prompt tuning, we identify a fundamental limitation in these methods: the Implicit Modality-Reduction bottleneck. By conditioning prompts solely on the observed modalities, they inadvertently restrict the reasoning scope of MTs to the modality-reduced subspace, cutting off access to the latent information sources of the missing modalities. To overcome this limitation, we propose AOEPT, which pioneers a novel modal-contextualized prompting fashion. Specifically, we introduce lightweight Modal-Contextualized Prompts (MCPs) that distill global modality-wise priors from training data, serving as latent repositories of the information sources for missing modalities. Conditioned on the remaining modalities, these MCPs are instantiated into instance-aware prompts that selectively augment missing-modality information for each sample, thereby restoring the reasoning scope of MTs beyond the observed-modality-only subspace. Experiments across various multimodal benchmarks and backbones confirm the strong performance of AOEPT, with minimal computational overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e4cad67-7267-4f92-83c2-24df3edc984aBuilds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesDouwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami et al.NeurIPS 2020 · 1,022 citations
- Prompt-aligned Gradient for Prompt TuningBeier Zhu, Yulei Niu, Yucheng Han, Yue Wu et al.ICCV 2023 · 475 citations
Related papers
- Retrieval-Augmented Dynamic Prompt Tuning for Incomplete Multimodal LearningJian Lang, Zhangtao Cheng, Ting Zhong, Fan ZhouAAAI 2025 · 20 citations
- REDEEMing Modality Information Loss: Retrieval-Guided Conditional Generation for Severely Modality Missing LearningJian Lang, Rongpei Hong, Zhangtao Cheng, Ting Zhong et al.KDD 2025 · 5 citations
- Multimodal Prompting with Missing Modalities for Visual RecognitionYi-Lun Lee, Yi-Hsuan Tsai, Wei-Chen Chiu, Chen-Yu LeeCVPR 2023
- Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion RecognitionZirun Guo, Tao Jin, Zhou ZhaoACL 2024 · 33 citations
- Deep Correlated Prompting for Visual Recognition with Missing ModalitiesLianyu Hu, Tongkai Shi, Wei Feng, Fanhua Shang et al.NeurIPS 2024 · 37 citations
