CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models
Sreyan Ghosh, Ashish Seth, Sonal Kumar, Utkarsh Tyagi, Chandra Kiran Reddy Evuru, Ramaneswaran S., Sakshi Singh, Oriol Nieto, Ramani Duraiswami, Dinesh Manocha
摘要
A fundamental characteristic of audio is its compositional nature. Audio-language models (ALMs) trained using a contrastive approach (e.g., CLAP) that learns a shared representation between audio and language modalities have improved performance in many downstream applications, including zero-shot audio classification, audio retrieval, etc. However, the ability of these models to effectively perform compositional reasoning remains largely unexplored and necessitates additional research. In this paper, we propose CompA, a collection of two expert-annotated benchmarks with a majority of real-world audio samples, to evaluate compositional reasoning in ALMs. Our proposed CompA-order evaluates how well an ALM understands the order or occurrence of acoustic events in audio, and CompA-attribute evaluates attribute-binding of acoustic events. An instance from either benchmark consists of two audio-caption pairs, where both audios have the same acoustic events but with different compositions. An ALM is evaluated on how well it matches the right audio to the right caption. Using this benchmark, we first show that current ALMs perform only marginally better than random chance, thereby struggling with compositional reasoning. Next, we propose CompA-CLAP, where we fine-tune CLAP using a novel learning method to improve its compositional reasoning abilities. To train CompA-CLAP, we first propose improvements to contrastive training with composition-aware hard negatives, allowing for more focused training. Next, we propose a novel modular contrastive loss that helps the model learn fine-grained compositional understanding and overcomes the acute scarcity of openly available compositional audios. CompA-CLAP significantly improves over all our baseline models on the CompA benchmark, indicating its superior compositional reasoning capabilities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language ModelsSreyan Ghosh, Arushi Goel, Jaehyeon Kim, Sonal Kumar 等NeurIPS 2025 · 被引用 299 次
- OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLMHanrong Ye, Chao-Han Huck Yang, Arushi Goel, Wei Huang 等ICLR 2026 · 被引用 64 次
- ATRI: Mitigating Multilingual Audio Text Retrieval Inconsistencies by Reducing Data Distribution ErrorsYuguo Yin, Yuxin Xie, Wenyuan Yang, Dongchao Yang 等ACL 2025 · 被引用 12 次
- Advancing Multi-grained Alignment for Contrastive Language-Audio Pre-trainingYiming Li, Zhifang Guo, Xiangdong Wang, Hong LiuACM MM 2024 · 被引用 9 次
- CCFQA: A Benchmark for Cross-Lingual and Cross-Modal Speech and Text Factuality EvaluationYexing Du, Kaiyuan Liu, Youcheng Pan, Zheng Chu 等AAAI 2026 · 被引用 4 次
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei 等ICML 2023 · 被引用 773 次
- Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion ModelsRongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren 等ICML 2023 · 被引用 469 次
相关 Paper
- Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional UnderstandingLe Zhang, Rabiul Awal, Aishwarya AgrawalCVPR 2024 · 被引用 7 次
- No Hard Negatives Required: Concept Centric Learning Leads to Compositionality without Degrading Zero-shot Capabilities of Contrastive ModelsHai X. Pham, David T. Hoffmann, Ricardo Guerrero, Brais MartínezCVPR 2026 · 被引用 1 次
- Interpretable Composition Attribution Enhancement for Visio-linguistic Compositional UnderstandingWei Li, Zhen Huang, Xinmei Tian, Le Lu 等EMNLP 2024
- When and Why Vision-Language Models Behave like Bags-Of-Words, and What to Do About It?Mert Yüksekgönül, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky 等ICLR 2023 · 被引用 37 次
- CoLLAT: On Adding Fine-grained Audio Understanding to Language Models using Token-Level Locked-Language TuningDadallage A. R. Silva, Spencer Whitehead, Christopher T. Lengerich, Hugh LeatherNeurIPS 2023 · 被引用 11 次
