Lune

KDD2026顶会

LSAR: Sparse Lexical Representation Learning for Efficient and Interpretable Audio Retrieval

Haoyue Li, Yuzhe Bai, Li Niu

2026年份

摘要

As Multimodal Large Language Models (MLLMs) expand the scope of retrieval-augmented generation, recommendation, and multimedia search, audio retrieval is expected to become a dependable retrieval component. Yet existing systems struggle to reconcile lexical precision, non-verbal acoustic evidence, and efficient, transparent retrieval. Current approaches mainly follow two paradigms. Cascaded pipelines transcribe audio with automatic speech recognition (ASR) and retrieve over text. They inherit the strengths of lexical matching, but amplify recognition errors and systematically discard non-verbal cues such as acoustic events and speaking style. Dense retrievers bypass transcription, yet they compress each clip into a single embedding, making relevance difficult to inspect and incurring non-trivial indexing and query-time overhead as collections grow. To address these limitations, we propose LSAR (Learned Sparse Audio Retrieval), the first learned sparse retrieval framework for audio that maps audio directly into a sparse lexical space compatible with inverted-index search. LSAR employs cross-modal sparse alignment and multiple complementary branches to model spoken content, acoustic context, and logic associations, producing an indexable representation with explicit term activations. At inference, retrieval proceeds directly from audio without an ASR transcription stage. Experiments on diverse benchmarks covering speech, captioning, and audio question answering show that LSAR effectively maps continuous audio representations into a sparse lexical space. Acting as a robust first-stage retriever, it enables high-recall candidate retrieval with low latency to significantly narrow the search space, while offering keyword-level interpretability. Overall, LSAR introduces a new retrieval paradigm for audio, establishing an interpretable, index-friendly primitive for real-time multimodal RAG and beyond.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖