AVQACL: A Novel Benchmark for Audio-Visual Question Answering Continual Learning
Kaixuan Wu, Xinde Li, Xinling Li, Chuanfei Hu, Guoliang Wu
摘要
In this paper, a novel benchmark for audio-visual question answering continual learning (AVQACL) is introduced, aiming to study fine-grained scene understanding and spatial-temporal reasoning in videos under a continual learning setting. To facilitate this multimodal continual leaning task, we create two audio-visual question answering continual learning datasets, named Split-AVQA and Split-MUSIC-AVQA based on the AVQA and MUSIC-AVQA datasets, respectively. The experimental results suggest that the model exhibits limited cognitive and reasoning abilities and experiences catastrophic forgetting when processing three modalities simultaneously in a continuous data stream. To address above challenges, we propose a novel continual learning method that incorporates question-guided cross-modal information fusion (QCIF) to focus on question-relevant details for improved feature representation and task-specific knowledge distillation with spatial-temporal feature constraints (TKD-STFC) to preserve the spatial-temporal reasoning knowledge acquired from previous dynamic scenarios. Furthermore, a question semantic consistency constraint (QSCC) is employed to ensure that the model maintains a consistent understanding of question semantics across tasks throughout the continual learning process. Extensive experimental results on Split-AVQA and Split-MUSIC-AVQA datasets illustrate that our method achieves state-of-the-art audio-visual question answering continual learning performance. The code is available at https://github.com/kx-wu/AVQACL .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper17
- SS-IL: Separated Softmax for Incremental LearningHongjoon Ahn, Jihwan Kwak, Subin Lim, Hyeonsu Bang 等ICCV 2021 · 被引用 209 次
- Learning to Answer Questions in Dynamic Audio-Visual ScenariosGuangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu 等CVPR 2022 · 被引用 101 次
- AVQA: A Dataset for Audio-Visual Question Answering on VideosPinci Yang, Xin Wang, Xuguang Duan, Hong Chen 等ACM MM 2022 · 被引用 60 次
- Cross-modal Prompts: Adapting Large Pre-trained Models for Audio-Visual Downstream TasksHaoyi Duan, Yan Xia, Mingze Zhou, Li Tang 等NeurIPS 2023 · 被引用 59 次
- Audio-Visual Class-Incremental LearningWeiguo Pian, Shentong Mo, Yunhui Guo, Yapeng TianICCV 2023 · 被引用 44 次
相关 Paper
- VQACL: A Novel Visual Question Answering Continual Learning SettingXi Zhang, Feifei Zhang, Changsheng XuCVPR 2023
- Re-evaluating Continual VQA: Toward Fair and Robust Evaluation for Multimodal Continual LearningZijian Gao, Zicheng Sun, Xingxing Zhang, Kele Xu 等CVPR 2026
- Continual Audio-Visual Sound SeparationWeiguo Pian, Yiyang Nan, Shijian Deng, Shentong Mo 等NeurIPS 2024 · 被引用 11 次
- Audio-Visual Adaptive Fusion Network for Question Answering Based on Contrastive LearningXujian Zhao, Yixin Wang, Peiquan JinAAAI 2025 · 被引用 4 次
- Query-Guided Spatial-Temporal-Frequency Interaction for Music Audio-Visual Question AnsweringKun Li, Michael Ying Yang, Sami Sebastian BrandtICLR 2026 · 被引用 1 次
