Query-Guided Spatial-Temporal-Frequency Interaction for Music Audio-Visual Question Answering
Kun Li, Michael Ying Yang, Sami Sebastian Brandt
Abstract
Audio--Visual Question Answering (AVQA) is a challenging multimodal task that requires jointly reasoning over audio, visual, and textual information in a given video to answer natural language questions. Inspired by recent advances in Video QA, many existing AVQA approaches primarily focus on visual information processing, leveraging pre-trained models to extract object-level and motion-level representations. However, in those methods, the audio input is primarily treated as complementary to video analysis, and the textual question information contributes minimally to audio--visual understanding, as it is typically integrated only in the final stages of reasoning. To address these limitations, we propose a novel Query-guided Spatial--Temporal--Frequency (QSTar) interaction method, which effectively incorporates question-guided clues and exploits the distinctive frequency-domain characteristics of audio signals, alongside spatial and temporal perception, to enhance audio--visual understanding. Furthermore, we introduce a Query Context Reasoning (QCR) block inspired by prompting, which guides the model to focus more precisely on semantically relevant audio and visual features. Extensive experiments conducted on several AVQA benchmarks demonstrate the effectiveness of our proposed method, achieving significant performance improvements over existing Audio QA, Visual QA, Video QA, and AVQA approaches. The code and pretrained models will be released after publication.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Learning to Answer Questions in Dynamic Audio-Visual ScenariosGuangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu et al.CVPR 2022 · 101 citations
- Token Merging: Your ViT But FasterDaniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang et al.ICLR 2023 · 62 citations
- AVQA: A Dataset for Audio-Visual Question Answering on VideosPinci Yang, Xin Wang, Xuguang Duan, Hong Chen et al.ACM MM 2022 · 60 citations
- Cross-modal Prompts: Adapting Large Pre-trained Models for Audio-Visual Downstream TasksHaoyi Duan, Yan Xia, Mingze Zhou, Li Tang et al.NeurIPS 2023 · 59 citations
Related papers
- Boosting Audio Visual Question Answering via Key Semantic-Aware CuesGuangyao Li, Henghui Du, Di HuACM MM 2024 · 16 citations
- Progressive Spatio-temporal Perception for Audio-Visual Question AnsweringGuangyao Li, Wenxuan Hou, Di HuACM MM 2023 · 39 citations
- Object-Aware Adaptive-Positivity Learning for Audio-Visual Question AnsweringZhangbin Li, Dan Guo, Jinxing Zhou, Jing Zhang et al.AAAI 2024 · 30 citations
- AVQACL: A Novel Benchmark for Audio-Visual Question Answering Continual LearningKaixuan Wu, Xinde Li, Xinling Li, Chuanfei Hu et al.CVPR 2025
- Audio-Visual Adaptive Fusion Network for Question Answering Based on Contrastive LearningXujian Zhao, Yixin Wang, Peiquan JinAAAI 2025 · 4 citations
