Discrepancy-Aware Attention Network for Enhanced Audio-Visual Generalized Zero-Shot Learning
Runlin Yu, Yipu Gong, Wenrui Li, Aiwen Sun, Mengren Zheng
摘要
Audio-visual Generalized Zero-Shot Learning ((G)ZSL) has attracted significant attention for its ability to identify unseen classes in general video classification tasks. However, modality imbalance in (G)ZSL leads to over-reliance on the optimal modality, reducing discriminative capabilities for unseen classes. Though recent studies have attempted to address this issue, two challenges still remain unsolved: (a) Quality discrepancies, where modalities offer differing quantities and qualities of information for the same concept. (b) Content discrepancies, where the contributions of different samples within the same modality exhibit significant differences. To address these challenges, we propose a Discrepancy-Aware Attention Network (DAAN) for Enhanced Audio-Visual (G)ZSL. Our approach introduces a Redundant-Noise Mitigation Attention (RNMA) unit to minimize content discrepancies by mitigating redundant information in modalities and a Contrastive Sample Gradient Modulation (CSGM) mechanism to adjust gradient magnitudes and balance quality discrepancies. We quantify modality contributions by integrating optimization and convergence rate for more precise gradient modulation in CSGM. Experiments demonstrate DAAN achieves state-of-the-art performance on benchmark datasets, with ablation studies validating the effectiveness of individual modules. Code is available at https://github.com/xiaoxinning/DAAN-GZSL.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Audiovisual Generalised Zero-shot Learning with Cross-modal Attention and LanguageOtniel-Bogdan Mercea, Lukas Riesch, A. Sophia Koepke, Zeynep AkataCVPR 2022 · 被引用 54 次
- Balancing Cross-Modal Attention for Generalized Zero-Shot LearningZhijie Rao, Jingcai GuoACM MM 2025
- A Variational Autoencoder with Deep Embedding Model for Generalized Zero-Shot LearningPeirong Ma, Xiao HuAAAI 2020 · 被引用 43 次
- Progressive Semantic-Visual Mutual Adaption for Generalized Zero-Shot LearningMan Liu, Feng Li, Chunjie Zhang, Yunchao Wei 等CVPR 2023
- Semantics Disentangling for Generalized Zero-Shot LearningZhi Chen, Yadan Luo, Ruihong Qiu, Sen Wang 等ICCV 2021 · 被引用 143 次
