Towards Unified Vision-Language Models with Incomplete Multi-Modal Inputs
Xiang Fang, Wanlong Fang, Changshuo Wang, Keke Tang, Daizong Liu, Siyi Wang, Wei Ji
Abstract
Video-Language Models (VLMs) have demonstrated impressive multi-modal reasoning capabilities across diverse computer vision applications. However, these VLMs are taskspecific and assume that both video and language inputs are complete. However, real-world VLM applications might face challenges due to deactivated sensors (e.g., cameras are unavailable due to data privacy), yielding modality-incomplete data and leading to inconsistency between training and testing data. While straightforward incomplete input can boast training generalization-ability and lead to training failure, its potential risks to VLMs regarding safety and trustworthiness have been largely neglected. To this end, we make the first attempt to propose a unified incomplete video-language model to process the incomplete multi-modal inputs. Extensive experimental results show that our method can serve as a plugand-play module for previous works to improve their performance in various multi-modal tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4627ac36-b4d0-47de-bc8c-c5dabccb6806Cited by top-tier papers8
- Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action PoliciesZhixuan Liang, Yizhuo Li, Tianshuo Yang, CHENGYUE WU et al.ICML 2026 · 86 citations
- Immuno-VLM: Immunizing Large Vision-Language Models via Generative Semantic Antibodies for Open-World TrustworthinessXiang Fang, Wanlong Fang, Wei JiICML 2026 · 17 citations
- CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric ReasoningXiang Fang, Wanlong Fang, Changshuo WangCVPR 2026 · 17 citations
- Not All Inputs Are Valid: Towards Open-Set Video Moment Retrieval using LanguageXiang Fang, Wanlong Fang, Daizong Liu, Xiaoye Qu et al.ACM MM 2024 · 8 citations
- CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent AttentionZhongzhu Zhou, Fengxiang Bie, Ziyan Chen, Zhenyu Zhang et al.ICLR 2026 · 4 citations
Builds on30
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- What Makes Multi-Modal Learning Better than Single (Provably)Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen et al.NeurIPS 2021 · 404 citations
- Self-Chained Image-Language Model for Video Localization and Question AnsweringShoubin Yu, Jaemin Cho, Prateek Yadav, Mohit BansalNeurIPS 2023 · 281 citations
- X-Pool: Cross-Modal Language-Video Attention for Text-Video RetrievalSatya Krishna Gorti, Noël Vouitsis, Junwei Ma, Keyvan Golestan et al.CVPR 2022 · 190 citations
Related papers
- Learning Unseen Modality InteractionYunhua Zhang, Hazel Doughty, Cees SnoekNeurIPS 2023 · 16 citations
- Miss-ReID: Delivering Robust Multi-Modality Object Re-Identification Despite Missing ModalitiesRuida XiNeurIPS 2025 · 4 citations
- Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry PriorsDuo Zheng, Shijia Huang, Yanyang Li, Liwei WangNeurIPS 2025 · 130 citations
- Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional TokenizationYang Jin, Zhicheng Sun, Kun Xu, Kun Xu et al.ICML 2024 · 94 citations
- LIMSSR: LLM-Driven Sequence-to-Score Reasoning under Training-Time Incomplete Multimodal ObservationsHuangbiao Xu, huanqi wu, Xiao Ke, Yuxin PengICML 2026 · 1 citation
