Learning to Answer Questions in Dynamic Audio-Visual Scenarios
Guangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu, Ji-Rong Wen, Di Hu
Abstract
In this paper, we focus on the Audio-Visual Question Answering (AVQA) task, which aims to answer questions regarding different visual objects, sounds, and their associations in videos. The problem requires comprehensive multimodal understanding and spatio-temporal reasoning over audio-visual scenes. To benchmark this task and facilitate our study, we introduce a large-scale MUSIC-AVQA dataset, which contains more than 45K question-answer pairs covering 33 different question templates spanning over different modalities and question types. We develop several baselines and introduce a spatio-temporal grounded audio-visual network for the AVQA problem. Our results demonstrate that AVQA benefits from multisensory perception and our model outperforms recent A-, V-, and AVQA approaches. We believe that our built dataset has the potential to serve as testbed for evaluating and promoting progress in audio-visual scene understanding and spatio-temporal reasoning. Code and dataset: http://gewu-lab.github.io/MUSIC-AVQA/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 48cbb05b-2fea-4b6e-93d8-2c742a1100e2Cited by top-tier papers93
- Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language ModelsSreyan Ghosh, Arushi Goel, Jaehyeon Kim, Sonal Kumar et al.NeurIPS 2025 · 299 citations
- VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and DatasetSihan Chen, Handong Li, Qunbo Wang, Zijia Zhao et al.NeurIPS 2023 · 246 citations
- Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue AbilitiesZhifeng Kong, Arushi Goel, Rohan Badlani, Wei Ping et al.ICML 2024 · 207 citations
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMsJack Hong, Shilin Yan, Jiayin Cai, Xiaolong Jiang et al.ICLR 2026 · 162 citations
- Achieving Cross Modal Generalization with Multimodal Unified RepresentationYan Xia, Hai Huang, Jieming Zhu, Zhou ZhaoNeurIPS 2023 · 84 citations
Builds on18
- Dual Attention Matching for Audio-Visual Event LocalizationYu Wu, Linchao Zhu, Yan Yan, Yi YangICCV 2019 · 233 citations
- MultiModalQA: complex question answering over text, tables and imagesAlon Talmor, Ori Yoran, Amnon Catav, Dan Lahav et al.ICLR 2021 · 229 citations
- Self-supervised Video Object Segmentation by Motion GroupingCharig Yang, Hala Lamdouar, Erika Lu, Andrew Zisserman et al.ICCV 2021 · 188 citations
- Discriminative Sounding Objects Localization via Self-supervised Audiovisual MatchingDi Hu, Rui Qian, Minyue Jiang, Xiao Tan et al.NeurIPS 2020 · 156 citations
- Recursive Visual Sound Separation Using Minus-Plus NetXudong Xu, Bo Dai, Dahua LinICCV 2019 · 95 citations
Related papers
- AVQA: A Dataset for Audio-Visual Question Answering on VideosPinci Yang, Xin Wang, Xuguang Duan, Hong Chen et al.ACM MM 2022 · 60 citations
- Progressive Spatio-temporal Perception for Audio-Visual Question AnsweringGuangyao Li, Wenxuan Hou, Di HuACM MM 2023 · 39 citations
- Pano-AVQA: Grounded Audio-Visual Question Answering on 360° VideosHeeseung Yun, Youngjae Yu, Wonsuk Yang, Kangil Lee et al.ICCV 2021 · 124 citations
- Boosting Audio Visual Question Answering via Key Semantic-Aware CuesGuangyao Li, Henghui Du, Di HuACM MM 2024 · 16 citations
- Object-Aware Adaptive-Positivity Learning for Audio-Visual Question AnsweringZhangbin Li, Dan Guo, Jinxing Zhou, Jing Zhang et al.AAAI 2024 · 30 citations
