Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval
Haejun Yoo, Yongseop Shin, Insung Lee, Myoung-Wan Koo, Du-Seong Chang
Abstract
Audio-text retrieval systems based on Contrastive Language-Audio Pretraining (CLAP) achieve strong performance on traditional benchmarks; however, these benchmarks rely on caption-style queries that differ substantially from real-world search behavior, limiting their assessment of practical retrieval robustness. We present Omni-Embed-Audio (OEA), a retrievaloriented encoder leveraging multimodal LLMs with native audio understanding. To systematically evaluate robustness beyond captionstyle queries, we introduce User-Intent Queries (UIQs)-five formulations reflecting natural search behaviors: questions, commands, keyword tags, paraphrases, and exclusion-based negative queries. For negative queries, we develop a hard negative mining pipeline and propose discrimination metrics (HNSR, TFR) assessing models' ability to suppress acoustically similar distractors. Experiments on Au-dioCaps, Clotho, and MECAT show that OEA achieves comparable text-to-audio retrieval performance to state-of-the-art M2D-CLAP, while demonstrating clear advantages in two critical areas: (1) dominant text-to-text retrieval (+22% relative improvement), and (2) substantially superior hard negative discrimination (+4.3%p HNSR@10, +34.7% relative TFR@10). Mechanism ablations attribute these discrimination gains chiefly to OEA's audio embeddings, which place confusable clips significantly farther apart; the exclusion signal itself rides on query word order-an order sensitivity that cross-model controls show OEA shares with CLAP text encoders rather than uniquely possessing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 21558f25-abb4-4b00-8bcd-573112fac85bBuilds on1
Related papers
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei et al.ICML 2023 · 773 citations
- ATRI: Mitigating Multilingual Audio Text Retrieval Inconsistencies by Reducing Data Distribution ErrorsYuguo Yin, Yuxin Xie, Wenyuan Yang, Dongchao Yang et al.ACL 2025 · 12 citations
- FIGMA: Towards FIne-Grained Music retrievAlNishit Anand, Ashish Seth, Sreyan Ghosh, Dinesh Manocha et al.ACL 2026
- CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled VideosHao-Wen Dong, Naoya Takahashi, Yuki Mitsufuji, Julian J. McAuley et al.ICLR 2023 · 3 citations
- MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding TasksYadong Niu, TIANZI WANG, Heinrich Dinkel, Xingwei Sun et al.ICML 2026 · 11 citations
