Learning Spatially-Aware Language and Audio Embeddings
Bhavika Devnani, Skyler Seto, Zakaria Aldeneh, Alessandro Toso, Elena Menyaylenko, Barry-John Theobald, Jonathan Sheaffer, Miguel Sarabia
Abstract
Humans can picture a sound scene given an imprecise natural language description. For example, it is easy to imagine an acoustic environment given a phrase like"the lion roar came from right behind me!". For a machine to have the same degree of comprehension, the machine must know what a lion is (semantic attribute), what the concept of"behind"is (spatial attribute) and how these pieces of linguistic information align with the semantic and spatial attributes of the sound (what a roar sounds like when its coming from behind). State-of-the-art audio foundation models which learn to map between audio scenes and natural textual descriptions, are trained on non-spatial audio and text pairs, and hence lack spatial awareness. In contrast, sound event localization and detection models are limited to recognizing sounds from a fixed number of classes, and they localize the source to absolute position (e.g., 0.2m) rather than a position described using natural language (e.g.,"next to me"). To address these gaps, we present ELSA a spatially aware-audio and text embedding model trained using multimodal contrastive learning. ELSA supports non-spatial audio, spatial audio, and open vocabulary text captions describing both the spatial and semantic components of sound. To train ELSA: (a) we spatially augment the audio and captions of three open-source audio datasets totaling 4,738 hours of audio, and (b) we design an encoder to capture the semantics of non-spatial audio, and the semantics and spatial attributes of spatial audio using contrastive learning. ELSA is competitive with state-of-the-art for both semantic retrieval and 3D source localization. In particular, ELSA achieves +2.8% mean audio-to-text and text-to-audio R@1 above the baseline, and outperforms by -11.6 mean-absolute-error in 3D source localization over the baseline.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 118331df-6d71-4e5c-b022-1629b04d314aCited by top-tier papers5
- SPHERE: Unveiling Spatial Blind Spots in Vision-Language Models Through Hierarchical EvaluationWenyu Zhang, Wei En Ng, Lixin Ma, Yuwen Wang et al.ACL 2025 · 20 citations
- PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMsArtem Dementyev, Wazeer Zulfikar, Sinan Hersek, Pascal Getreuer et al.ICML 2026 · 4 citations
- Hear you are: Teaching LLMs Spatial Reasoning with Vision and Spatial SoundHyeonggon Ryu, Joon Son Chung, David HarwathCVPR 2026 · 4 citations
- SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video GenerationKien T. Pham, Yingqing He, Yazhou Xing, Qifeng Chen et al.ACM MM 2025 · 1 citation
- JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical EnvironmentsZhan Liu, Changli Tang, Yuxin WANG, Zhiyuan Zhu et al.ICML 2026
Builds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- SALMONN: Towards Generic Hearing Abilities for Large Language ModelsChangli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen et al.ICLR 2024 · 557 citations
- How Much Can CLIP Benefit Vision-and-Language Tasks?Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal et al.ICLR 2022 · 503 citations
- Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion ModelsRongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren et al.ICML 2023 · 469 citations
Related papers
- Improving Sound Source Localization with Joint Slot Attention on Image and AudioInho Kim, Youngkil Song, Jicheol Park, Won Hwa Kim et al.CVPR 2025
- Object-aware Sound Source Localization via Audio-Visual Scene UnderstandingSung Jin Um, Dongjin Kim, Sangmin Lee, Jung Uk KimCVPR 2025
- FLAM: Frame-Wise Language-Audio ModelingYusong Wu, Christos Tsirigotis, Ke Chen, Cheng-Zhi Anna Huang et al.ICML 2025
- Gotta Hear Them All: Towards Sound Source Aware Audio GenerationWei Guo, Heng Wang, Jianbo Ma, Weidong CaiAAAI 2026 · 2 citations
- SoundingActions: Learning How Actions Sound from Narrated Egocentric VideosChangan Chen, Kumar Ashutosh, Rohit Girdhar, David Harwath et al.CVPR 2024
