Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval
Haejun Yoo, Yongseop Shin, Insung Lee, Myoung-Wan Koo, Du-Seong Chang
摘要
Audio-text retrieval systems based on Contrastive Language-Audio Pretraining (CLAP) achieve strong performance on traditional benchmarks; however, these benchmarks rely on caption-style queries that differ substantially from real-world search behavior, limiting their assessment of practical retrieval robustness. We present Omni-Embed-Audio (OEA), a retrievaloriented encoder leveraging multimodal LLMs with native audio understanding. To systematically evaluate robustness beyond captionstyle queries, we introduce User-Intent Queries (UIQs)-five formulations reflecting natural search behaviors: questions, commands, keyword tags, paraphrases, and exclusion-based negative queries. For negative queries, we develop a hard negative mining pipeline and propose discrimination metrics (HNSR, TFR) assessing models' ability to suppress acoustically similar distractors. Experiments on Au-dioCaps, Clotho, and MECAT show that OEA achieves comparable text-to-audio retrieval performance to state-of-the-art M2D-CLAP, while demonstrating clear advantages in two critical areas: (1) dominant text-to-text retrieval (+22% relative improvement), and (2) substantially superior hard negative discrimination (+4.3%p HNSR@10, +34.7% relative TFR@10). Mechanism ablations attribute these discrimination gains chiefly to OEA's audio embeddings, which place confusable clips significantly farther apart; the exclusion signal itself rides on query word order-an order sensitivity that cross-model controls show OEA shares with CLAP text encoders rather than uniquely possessing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei 等ICML 2023 · 被引用 773 次
- ATRI: Mitigating Multilingual Audio Text Retrieval Inconsistencies by Reducing Data Distribution ErrorsYuguo Yin, Yuxin Xie, Wenyuan Yang, Dongchao Yang 等ACL 2025 · 被引用 12 次
- FIGMA: Towards FIne-Grained Music retrievAlNishit Anand, Ashish Seth, Sreyan Ghosh, Dinesh Manocha 等ACL 2026
- CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled VideosHao-Wen Dong, Naoya Takahashi, Yuki Mitsufuji, Julian J. McAuley 等ICLR 2023 · 被引用 3 次
- MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding TasksYadong Niu, TIANZI WANG, Heinrich Dinkel, Xingwei Sun 等ICML 2026 · 被引用 11 次
