Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question Answering
Jie Ma, Min Hu, Pinghui Wang, Wangchun Sun, Lingyun Song, Hongbin Pei, Jun Liu, Youtian Du
摘要
Audio-Visual Question Answering (AVQA) is a complex multi-modal reasoning task, demanding intelligent systems to accurately respond to natural language queries based on audio-video input pairs. Nevertheless, prevalent AVQA approaches are prone to overlearning dataset biases, resulting in poor robustness. Furthermore, current datasets may not provide a precise diagnostic for these methods. To tackle these challenges, firstly, we propose a novel dataset, MUSIC-AVQA-R, crafted in two steps: rephrasing questions within the test split of a public dataset (MUSIC-AVQA) and subsequently introducing distribution shifts to split questions. The former leads to a large, diverse test space, while the latter results in a comprehensive robustness evaluation on rare, frequent, and overall questions. Secondly, we propose a robust architecture that utilizes a multifaceted cycle collaborative debiasing strategy to overcome bias learning. Experimental results show that this architecture achieves state-of-the-art performance on MUSIC-AVQA-R, notably obtaining a significant improvement of 9.32%. Extensive ablation experiments are conducted on the two datasets mentioned to analyze the component effectiveness within the debiasing strategy. Additionally, we highlight the limited robustness of existing multi-modal QA methods through the evaluation on our dataset. We also conduct experiments combining various baselines with our proposed strategy on two datasets to verify its plug-and-play capability. Our dataset and code are available at https://github.com/reml-group/MUSIC-AVQA-R.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Progressive Homeostatic and Plastic Prompt Tuning for Audio-Visual Multi-Task Incremental LearningJiong Yin, Liang Li, Jiehua Zhang, Yuhan Gao 等ICCV 2025 · 被引用 3 次
- EgoAVU: Egocentric Audio-Visual UnderstandingAshish Seth, Xinhao Mei, Changsheng Zhao, Varun Nagaraja 等CVPR 2026 · 被引用 1 次
- Query-Guided Spatial-Temporal-Frequency Interaction for Music Audio-Visual Question AnsweringKun Li, Michael Ying Yang, Sami Sebastian BrandtICLR 2026 · 被引用 1 次
- Question-Aware Gaussian Experts for Audio-Visual Question AnsweringHongyeob Kim, Inyoung Jung, Dayoon Suh, Youjia Zhang 等CVPR 2025
- VGGSounder: Audio-Visual Evaluations for Foundation ModelsDaniil Zverev, Thaddäus Wiedemer, Ameya Prabhu, Matthias Bethge 等ICCV 2025
它引用的顶会 Paper30
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao 等CVPR 2022 · 被引用 2,138 次
- Large Language Models Are Not Robust Multiple Choice SelectorsChujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou 等ICLR 2024 · 被引用 424 次
- Zero-Shot Video Question Answering via Frozen Bidirectional Language ModelsAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev 等NeurIPS 2022 · 被引用 305 次
- The Effect of Natural Distribution Shift on Question Answering ModelsJohn Miller, Karl Krauth, Benjamin Recht, Ludwig SchmidtICML 2020 · 被引用 158 次
- Pano-AVQA: Grounded Audio-Visual Question Answering on 360° VideosHeeseung Yun, Youngjae Yu, Wonsuk Yang, Kangil Lee 等ICCV 2021 · 被引用 124 次
相关 Paper
- COCA: COllaborative CAusal Regularization for Audio-Visual Question AnsweringMingrui Lao, Nan Pu, Yu Liu, Kai He 等AAAI 2023 · 被引用 28 次
- AVQA: A Dataset for Audio-Visual Question Answering on VideosPinci Yang, Xin Wang, Xuguang Duan, Hong Chen 等ACM MM 2022 · 被引用 60 次
- Learning to Answer Questions in Dynamic Audio-Visual ScenariosGuangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu 等CVPR 2022 · 被引用 101 次
- Object-Aware Adaptive-Positivity Learning for Audio-Visual Question AnsweringZhangbin Li, Dan Guo, Jinxing Zhou, Jing Zhang 等AAAI 2024 · 被引用 30 次
- Audio-Visual Adaptive Fusion Network for Question Answering Based on Contrastive LearningXujian Zhao, Yixin Wang, Peiquan JinAAAI 2025 · 被引用 4 次
