CoLLAT: On Adding Fine-grained Audio Understanding to Language Models using Token-Level Locked-Language Tuning
Dadallage A. R. Silva, Spencer Whitehead, Christopher T. Lengerich, Hugh Leather
Abstract
Humans can easily understand various audio concepts, but conventional audio classification models fail due to their inability to predict unseen classes during training. To address this challenge, recent literature has explored contrastive language-audio pretraining to learn an audio understanding model using natural language supervision from a pretrained language model. However, despite their reasonable zero-shot performance in audio understanding, these models typically fail to achieve optimal performance while preserving the text understanding capabilities of the pretrained language model. They also perform poorly when comprehending audio clips with multiple audio concepts. To bridge these gaps, we propose CoLLAT :
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ccf2e8d5-6539-4c25-979c-4cfe73700a8fCited by top-tier papers3
- ATRI: Mitigating Multilingual Audio Text Retrieval Inconsistencies by Reducing Data Distribution ErrorsYuguo Yin, Yuxin Xie, Wenyuan Yang, Dongchao Yang et al.ACL 2025 · 12 citations
- SupCLAP: Controlling Optimization Trajectory Drift in Audio-Text Contrastive Learning with Support Vector RegularizationJiehui Luo, Yuguo Yin, Yuxin Xie, Jinghan Ru et al.ICLR 2026 · 3 citations
- M3Net: Efficient Time-Frequency Integration Network with Mirror Attention for Audio Classification on EdgeXuanming Jiang, Baoyi An, Guoshuai Zhao, Xueming QianAAAI 2025 · 2 citations
Builds on7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 1,214 citations
- TS2Vec: Towards Universal Representation of Time SeriesZhihan Yue, Yujing Wang, Juanyong Duan, Tianmeng Yang et al.AAAI 2022 · 938 citations
Related papers
- Advancing Multi-grained Alignment for Contrastive Language-Audio Pre-trainingYiming Li, Zhifang Guo, Xiangdong Wang, Hong LiuACM MM 2024 · 9 citations
- CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled VideosHao-Wen Dong, Naoya Takahashi, Yuki Mitsufuji, Julian J. McAuley et al.ICLR 2023 · 3 citations
- CompA: Addressing the Gap in Compositional Reasoning in Audio-Language ModelsSreyan Ghosh, Ashish Seth, Sonal Kumar, Utkarsh Tyagi et al.ICLR 2024 · 53 citations
- Listenable Maps for Zero-Shot Audio ClassifiersFrancesco Paissan, Luca Della Libera, Mirco Ravanelli, Cem SubakanNeurIPS 2024 · 5 citations
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei et al.ICML 2023 · 773 citations
