Audiovisual Generalised Zero-shot Learning with Cross-modal Attention and Language
Otniel-Bogdan Mercea, Lukas Riesch, A. Sophia Koepke, Zeynep Akata
摘要
Learning to classify video data from classes not included in the training data, i.e. video-based zero-shot learning, is challenging. We conjecture that the natural alignment between the audio and visual modalities in video data provides a rich training signal for learning discriminative multi-modal representations. Focusing on the relatively underexplored task of audio-visual zero-shot learning, we propose to learn multi-modal representations from audio- visual data using cross-modal attention and exploit textual label embeddings for transferring knowledge from seen classes to unseen classes. Taking this one step further, in our generalised audio-visual zero-shot learning setting, we include all the training classes in the test-time search space which act as distractors and increase the difficulty while making the setting more realistic. Due to the lack of a unified benchmark in this domain, we introduce a (generalised) zero-shot learning benchmark on three audio-visual datasets of varying sizes and difficulty, VGGSound, UCF, and ActivityNet, ensuring that the unseen test classes do not appear in the dataset used for supervised training of the backbone deep models. Comparing multiple relevant and recent methods, we demonstrate that our proposed AVCA model achieves state-of-the-art performance on all three datasets. Code and data are available at https://github.com/ExplainableML/AVCA-GZSL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Hyperbolic Audio-visual Zero-shot LearningJie Hong, Zeeshan Hayder, Junlin Han, Pengfei Fang 等ICCV 2023 · 被引用 27 次
- Curriculum-Listener: Consistency- and Complementarity-Aware Audio-Enhanced Temporal Sentence GroundingHoulun Chen, Xin Wang, Xiaohan Lan, Hong Chen 等ACM MM 2023 · 被引用 13 次
- Task-Driven Exploration: Decoupling and Inter-Task Feedback for Joint Moment Retrieval and Highlight DetectionJin Yang, Ping Wei, Huan Li, Ziyang RenCVPR 2024 · 被引用 13 次
- Audio Does Matter: Importance-Aware Multi-Granularity Fusion for Video Moment RetrievalJunan Lin, Daizong Liu, Xianke Chen, Xiaoye Qu 等ACM MM 2025 · 被引用 3 次
- AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMsSanjoy Chowdhury, Sayan Nag, Subhrajyoti Dasgupta, Yaoting Wang 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper17
- Attention Bottlenecks for Multimodal FusionArsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen 等NeurIPS 2021 · 被引用 884 次
- Self-Supervised Learning by Cross-Modal Audio-Video ClusteringHumam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani 等NeurIPS 2020 · 被引用 483 次
- Attribute Prototype Network for Zero-Shot LearningWenjia Xu, Yongqin Xian, Jiuniu Wang, Bernt Schiele 等NeurIPS 2020 · 被引用 392 次
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 被引用 271 次
- Co-Separating Sounds of Visual ObjectsRuohan Gao, Kristen GraumanICCV 2019 · 被引用 224 次
相关 Paper
- Generalized Zero-Shot Video Classification via Generative Adversarial NetworksMingyao Hong, Guorong Li, Xinfeng Zhang, Qingming HuangACM MM 2020 · 被引用 13 次
- A Variational Autoencoder with Deep Embedding Model for Generalized Zero-Shot LearningPeirong Ma, Xiao HuAAAI 2020 · 被引用 43 次
- Distilling Audio-Visual Knowledge by Compositional Contrastive LearningYanbei Chen, Yongqin Xian, A. Sophia Koepke, Ying Shan 等CVPR 2021
- Zero-shot Video Classification with Appropriate Web and Task Knowledge TransferJunbao Zhuo, Yan Zhu, Shuhao Cui, Shuhui Wang 等ACM MM 2022 · 被引用 11 次
- Discrepancy-Aware Attention Network for Enhanced Audio-Visual Generalized Zero-Shot LearningRunlin Yu, Yipu Gong, Wenrui Li, Aiwen Sun 等ACM MM 2025 · 被引用 1 次
