Retrieving to Recover: Towards Incomplete Audio-Visual Question Answering via Semantic-consistent Purification
Jiayu Zhang, Shuo Ye, Qilang Ye, Zihan Song, Jiajian Huang, Zitong Yu
摘要
Recent Audio-Visual Question Answering (AVQA) methods have advanced significantly. However, most AVQA methods lack effective mechanisms for handling missing modalities, suffering from severe performance degradation in real-world scenarios with data interruptions. Furthermore, prevailing methods for handling missing modalities predominantly rely on generative imputation to synthesize missing features. While partially effective, these methods tend to capture inter-modal commonalities but struggle to acquire unique, modality-specific knowledge within the missing data, leading to hallucinations and compromised reasoning accuracy. To tackle these challenges, we propose RScP, a novel framework that shifts the paradigm of missing modality handling from traditional generative imputation to retrieval-based recovery. Specifically, we leverage cross-modal retrieval via unified semantic embeddings to acquire missing domain-specific knowledge. To maximize semantic restoration, we introduce a context-aware adaptive purification mechanism that eliminates latent semantic noise within the retrieved data. Additionally, we employ a two-stage training strategy to explicitly model the semantic relationships between knowledge from different sources. Extensive experiments demonstrate that RScP significantly improves AVQA and enhances robustness in modal-incomplete scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Are Multimodal Transformers Robust to Missing Modality?Mengmeng Ma, Jian Ren, Long Zhao, Davide Testuggine 等CVPR 2022 · 被引用 153 次
- Learning to Answer Questions in Dynamic Audio-Visual ScenariosGuangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu 等CVPR 2022 · 被引用 101 次
- AVQA: A Dataset for Audio-Visual Question Answering on VideosPinci Yang, Xin Wang, Xuguang Duan, Hong Chen 等ACM MM 2022 · 被引用 60 次
- Progressive Spatio-temporal Perception for Audio-Visual Question AnsweringGuangyao Li, Wenxuan Hou, Di HuACM MM 2023 · 被引用 39 次
- Leveraging Knowledge of Modality Experts for Incomplete Multimodal LearningWenxin Xu, Hexin Jiang, Xuefeng LiangACM MM 2024 · 被引用 31 次
相关 Paper
- REDEEMing Modality Information Loss: Retrieval-Guided Conditional Generation for Severely Modality Missing LearningJian Lang, Rongpei Hong, Zhangtao Cheng, Ting Zhong 等KDD 2025 · 被引用 5 次
- MISSRAG: Addressing the Missing Modality Challenge in Multimodal Large Language ModelsVittorio Pipoli, Alessia Saporita, Federico Bolelli, Marcella Cornia 等ICCV 2025 · 被引用 4 次
- Retrieval-Augmented Dynamic Prompt Tuning for Incomplete Multimodal LearningJian Lang, Zhangtao Cheng, Ting Zhong, Fan ZhouAAAI 2025 · 被引用 20 次
- Causality-Aligned Semantic Recovery for Incomplete Cross-Modal RetrievalHaipeng Chen, Yu Liu, Xun Yang, Yuheng Liang 等AAAI 2026
- RAG4DMC: Retrieval-Augmented Generation for Data-Level Modality CompletionNingxin He, Yongheng Deng, Sheng Yue, Yongjian Fu 等ICLR 2026
