What Do Language Models Hear? Probing for Auditory Representations in Language Models
Jerry Ngo, Yoon Kim
Abstract
This work explores whether language models encode meaningfully grounded representations of sounds of objects. We learn a linear probe that retrieves the correct text representation of an object given a snippet of audio related to that object, where the sound representation is given by a pretrained audio model. This probe is trained via a contrastive loss that pushes the language representations and sound representations of an object to be close to one another. After training, the probe is tested on its ability to generalize to objects that were not seen during training. Across different language models and audio models, we find that the probe generalization is above chance in many cases, indicating that despite being trained only on raw text, language models encode grounded knowledge of sounds for some objects.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f4a2b758-82e1-4dfd-ade3-f4c73b308a68Cited by top-tier papers6
- Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter ItYulu Qin, Dheeraj Varghese, Adam Dahlgren Lindström, Lucia Donatelli et al.NeurIPS 2025 · 11 citations
- The Indra Representation Hypothesis for Multimodal AlignmentJianglin Lu, Hailing Wang, Kuo Yang, Yitian Zhang et al.NeurIPS 2025 · 8 citations
- Probing Neural Combinatorial Optimization ModelsZhiqin Zhang, Yining Ma, Zhiguang Cao, Hoong Chuin LauNeurIPS 2025 · 3 citations
- How Tokenization Limits Phonological Knowledge Representation in Language Models and How to Improve ThemDisen Liao, Freda ShiACL 2026 · 1 citation
- The Semantic Hub Hypothesis: Language Models Share Semantic Representations Across Languages and ModalitiesZhaofeng Wu, Xinyan Velocity Yu, Dani Yogatama, Jiasen Lu et al.ICLR 2025
Builds on8
- Climbing towards NLU: On Meaning, Form, and Understanding in the Age of DataEmily M. Bender, Alexander KollerACL 2020 · 914 citations
- Masked Autoencoders that ListenPo-Yao Huang, Hu Xu, Juncheng Li, Alexei Baevski et al.NeurIPS 2022 · 524 citations
- Language Models Represent Space and TimeWes Gurnee, Max TegmarkICLR 2024 · 303 citations
- Mapping Language Models to Grounded Conceptual SpacesRoma Patel, Ellie PavlickICLR 2022 · 197 citations
- Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic TaskKenneth Li, Aspen K. Hopkins, David Bau, Fernanda B. Viégas et al.ICLR 2023 · 60 citations
Related papers
- Revisiting Audio-language Pretraining for Learning General-purpose Audio RepresentationWei-Cheng Tseng, Xuanru Zhou, Mingyue Huo, Yiwen Shao et al.ACL 2026 · 2 citations
- CoLLAT: On Adding Fine-grained Audio Understanding to Language Models using Token-Level Locked-Language TuningDadallage A. R. Silva, Spencer Whitehead, Christopher T. Lengerich, Hugh LeatherNeurIPS 2023 · 11 citations
- Learning Spatially-Aware Language and Audio EmbeddingsBhavika Devnani, Skyler Seto, Zakaria Aldeneh, Alessandro Toso et al.NeurIPS 2024 · 31 citations
- CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled VideosHao-Wen Dong, Naoya Takahashi, Yuki Mitsufuji, Julian J. McAuley et al.ICLR 2023 · 3 citations
- AudioMosaic: Contrastive Masked Audio Representation LearningHanxun Huang, Qizhou Wang, Xingjun Ma, Cihang Xie et al.ICML 2026 · 2 citations
