Multi-grained Correspondence Learning of Audio-language Models for Few-shot Audio Recognition
Shengwei Zhao, Linhai Xu, Yuying Liu, Shaoyi Du
Abstract
Large-scale pre-trained audio-language models excel in general multi-modal representation, facilitating their adaptation to downstream audio recognition tasks in a data-efficient manner. However, existing few-shot audio recognition methods based on audio-language models primarily focus on learning coarse-grained correlations, which are not sufficient to capture the intricate matching patterns between the multi-level information of audio and the diverse characteristics of category concepts. To address this gap, we propose multi-grained correspondence learning for bootstrapping audio-language models to improve audio recognition with few training samples. This approach leverages generative models to enrich multi-modal representation learning, mining the multi-level information of audio alongside the diverse characteristics of category concepts. Multi-grained matching patterns are then established through multi-grained key-value cache and multi-grained cross-modal contrast, enhancing the alignment between audio and category concepts. Additionally, we incorporate optimal transport to tackle temporal misalignment and semantic intersection issues in fine-grained correspondence learning, enabling flexible fine-grained matching. Our method achieves state-of-the-art results on multiple benchmark datasets for few-shot audio recognition, with comprehensive ablation experiments validating its effectiveness.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Advancing Multi-grained Alignment for Contrastive Language-Audio Pre-trainingYiming Li, Zhifang Guo, Xiangdong Wang, Hong LiuACM MM 2024 · 9 citations
- AWT: Transferring Vision-Language Models via Augmentation, Weighting, and TransportationYuhan Zhu, Yuyang Ji, Zhiyu Zhao, Gangshan Wu et al.NeurIPS 2024 · 45 citations
- CoLLAT: On Adding Fine-grained Audio Understanding to Language Models using Token-Level Locked-Language TuningDadallage A. R. Silva, Spencer Whitehead, Christopher T. Lengerich, Hugh LeatherNeurIPS 2023 · 11 citations
- Twofold Debiasing Enhances Fine-Grained Learning with Coarse LabelsXin-yang Zhao, Jian Jin, Yangyang Li, Yazhou YaoAAAI 2025 · 2 citations
- TAPE: Task-Adaptive Prototype Evolution in Audio-Language Models for Fully Few-shot Class-incremental Audio ClassificationYunlong Gao, Wenxin Liang, Guanglu Wang, Senqi Guan et al.CVPR 2026
