Audiovisual Generalised Zero-shot Learning with Cross-modal Attention and Language
Otniel-Bogdan Mercea, Lukas Riesch, A. Sophia Koepke, Zeynep Akata
Abstract
Learning to classify video data from classes not included in the training data, i.e. video-based zero-shot learning, is challenging. We conjecture that the natural alignment between the audio and visual modalities in video data provides a rich training signal for learning discriminative multi-modal representations. Focusing on the relatively underexplored task of audio-visual zero-shot learning, we propose to learn multi-modal representations from audio- visual data using cross-modal attention and exploit textual label embeddings for transferring knowledge from seen classes to unseen classes. Taking this one step further, in our generalised audio-visual zero-shot learning setting, we include all the training classes in the test-time search space which act as distractors and increase the difficulty while making the setting more realistic. Due to the lack of a unified benchmark in this domain, we introduce a (generalised) zero-shot learning benchmark on three audio-visual datasets of varying sizes and difficulty, VGGSound, UCF, and ActivityNet, ensuring that the unseen test classes do not appear in the dataset used for supervised training of the backbone deep models. Comparing multiple relevant and recent methods, we demonstrate that our proposed AVCA model achieves state-of-the-art performance on all three datasets. Code and data are available at https://github.com/ExplainableML/AVCA-GZSL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 503b7020-5026-4da8-9482-39a5eb4e7d35Cited by top-tier papers8
- Hyperbolic Audio-visual Zero-shot LearningJie Hong, Zeeshan Hayder, Junlin Han, Pengfei Fang et al.ICCV 2023 · 27 citations
- Curriculum-Listener: Consistency- and Complementarity-Aware Audio-Enhanced Temporal Sentence GroundingHoulun Chen, Xin Wang, Xiaohan Lan, Hong Chen et al.ACM MM 2023 · 13 citations
- Task-Driven Exploration: Decoupling and Inter-Task Feedback for Joint Moment Retrieval and Highlight DetectionJin Yang, Ping Wei, Huan Li, Ziyang RenCVPR 2024 · 13 citations
- Audio Does Matter: Importance-Aware Multi-Granularity Fusion for Video Moment RetrievalJunan Lin, Daizong Liu, Xianke Chen, Xiaoye Qu et al.ACM MM 2025 · 3 citations
- AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMsSanjoy Chowdhury, Sayan Nag, Subhrajyoti Dasgupta, Yaoting Wang et al.ICCV 2025 · 1 citation
Builds on17
- Attention Bottlenecks for Multimodal FusionArsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen et al.NeurIPS 2021 · 884 citations
- Self-Supervised Learning by Cross-Modal Audio-Video ClusteringHumam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani et al.NeurIPS 2020 · 483 citations
- Attribute Prototype Network for Zero-Shot LearningWenjia Xu, Yongqin Xian, Jiuniu Wang, Bernt Schiele et al.NeurIPS 2020 · 392 citations
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 271 citations
- Co-Separating Sounds of Visual ObjectsRuohan Gao, Kristen GraumanICCV 2019 · 224 citations
Related papers
- Generalized Zero-Shot Video Classification via Generative Adversarial NetworksMingyao Hong, Guorong Li, Xinfeng Zhang, Qingming HuangACM MM 2020 · 13 citations
- A Variational Autoencoder with Deep Embedding Model for Generalized Zero-Shot LearningPeirong Ma, Xiao HuAAAI 2020 · 43 citations
- Distilling Audio-Visual Knowledge by Compositional Contrastive LearningYanbei Chen, Yongqin Xian, A. Sophia Koepke, Ying Shan et al.CVPR 2021
- Zero-shot Video Classification with Appropriate Web and Task Knowledge TransferJunbao Zhuo, Yan Zhu, Shuhao Cui, Shuhui Wang et al.ACM MM 2022 · 11 citations
- Discrepancy-Aware Attention Network for Enhanced Audio-Visual Generalized Zero-Shot LearningRunlin Yu, Yipu Gong, Wenrui Li, Aiwen Sun et al.ACM MM 2025 · 1 citation
