VGGSounder: Audio-Visual Evaluations for Foundation Models
Daniil Zverev, Thaddäus Wiedemer, Ameya Prabhu, Matthias Bethge, Wieland Brendel, A. Sophia Koepke
Abstract
The emergence of audio-visual foundation models underscores the importance of reliably assessing their multimodal understanding. The VGGSound dataset is commonly used as a benchmark for evaluating audio-visual classification. However, our analysis identifies several limitations of VGGSound, including incomplete labelling, partially overlapping classes, and misaligned modalities. These lead to distorted evaluations of auditory and visual capabilities. To address these limitations, we introduce VGGSounder, a comprehensively re-annotated, multi-label test set that extends VGGSound and is specifically designed to evaluate audio-visual foundation models. VGGSounder features detailed modality annotations, enabling precise analyses of modality-specific performance. Furthermore, we reveal model limitations by analysing performance degradation when adding another input modality with our new modality confusion metric. Our dataset and project page are available at https://vggsounder.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 81d647f1-6ef8-42f3-bdbd-a7f73b967e84Cited by top-tier papers1
Ask how each one uses itBuilds on32
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Attention Bottlenecks for Multimodal FusionArsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen et al.NeurIPS 2021 · 884 citations
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch et al.ICLR 2021 · 878 citations
- BEATs: Audio Pre-Training with Acoustic TokenizersSanyuan Chen, Yu Wu, Chengyi Wang, Shujie Liu et al.ICML 2023 · 568 citations
Related papers
- Audiovisual Generalised Zero-shot Learning with Cross-modal Attention and LanguageOtniel-Bogdan Mercea, Lukas Riesch, A. Sophia Koepke, Zeynep AkataCVPR 2022 · 54 citations
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMsJack Hong, Shilin Yan, Jiayin Cai, Xiaolong Jiang et al.ICLR 2026 · 162 citations
- EgoSound: Benchmarking Sound Understanding in Egocentric VideosBingwen Zhu, Yuqian Fu, Qiaole Dong, Guolei Sun et al.CVPR 2026 · 7 citations
- AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language ModelsSung-Bin Kim, Oh Hyun-Bin, JungMok Lee, Arda Senocak et al.ICLR 2025
- Gotta Hear Them All: Towards Sound Source Aware Audio GenerationWei Guo, Heng Wang, Jianbo Ma, Weidong CaiAAAI 2026 · 2 citations
