What Do Language Models Hear? Probing for Auditory Representations in Language Models
Jerry Ngo, Yoon Kim
摘要
This work explores whether language models encode meaningfully grounded representations of sounds of objects. We learn a linear probe that retrieves the correct text representation of an object given a snippet of audio related to that object, where the sound representation is given by a pretrained audio model. This probe is trained via a contrastive loss that pushes the language representations and sound representations of an object to be close to one another. After training, the probe is tested on its ability to generalize to objects that were not seen during training. Across different language models and audio models, we find that the probe generalization is above chance in many cases, indicating that despite being trained only on raw text, language models encode grounded knowledge of sounds for some objects.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter ItYulu Qin, Dheeraj Varghese, Adam Dahlgren Lindström, Lucia Donatelli 等NeurIPS 2025 · 被引用 11 次
- The Indra Representation Hypothesis for Multimodal AlignmentJianglin Lu, Hailing Wang, Kuo Yang, Yitian Zhang 等NeurIPS 2025 · 被引用 8 次
- Probing Neural Combinatorial Optimization ModelsZhiqin Zhang, Yining Ma, Zhiguang Cao, Hoong Chuin LauNeurIPS 2025 · 被引用 3 次
- How Tokenization Limits Phonological Knowledge Representation in Language Models and How to Improve ThemDisen Liao, Freda ShiACL 2026 · 被引用 1 次
- The Semantic Hub Hypothesis: Language Models Share Semantic Representations Across Languages and ModalitiesZhaofeng Wu, Xinyan Velocity Yu, Dani Yogatama, Jiasen Lu 等ICLR 2025
它引用的顶会 Paper8
- Climbing towards NLU: On Meaning, Form, and Understanding in the Age of DataEmily M. Bender, Alexander KollerACL 2020 · 被引用 914 次
- Masked Autoencoders that ListenPo-Yao Huang, Hu Xu, Juncheng Li, Alexei Baevski 等NeurIPS 2022 · 被引用 524 次
- Language Models Represent Space and TimeWes Gurnee, Max TegmarkICLR 2024 · 被引用 303 次
- Mapping Language Models to Grounded Conceptual SpacesRoma Patel, Ellie PavlickICLR 2022 · 被引用 197 次
- Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic TaskKenneth Li, Aspen K. Hopkins, David Bau, Fernanda B. Viégas 等ICLR 2023 · 被引用 60 次
相关 Paper
- Revisiting Audio-language Pretraining for Learning General-purpose Audio RepresentationWei-Cheng Tseng, Xuanru Zhou, Mingyue Huo, Yiwen Shao 等ACL 2026 · 被引用 2 次
- CoLLAT: On Adding Fine-grained Audio Understanding to Language Models using Token-Level Locked-Language TuningDadallage A. R. Silva, Spencer Whitehead, Christopher T. Lengerich, Hugh LeatherNeurIPS 2023 · 被引用 11 次
- Learning Spatially-Aware Language and Audio EmbeddingsBhavika Devnani, Skyler Seto, Zakaria Aldeneh, Alessandro Toso 等NeurIPS 2024 · 被引用 31 次
- CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled VideosHao-Wen Dong, Naoya Takahashi, Yuki Mitsufuji, Julian J. McAuley 等ICLR 2023 · 被引用 3 次
- AudioMosaic: Contrastive Masked Audio Representation LearningHanxun Huang, Qizhou Wang, Xingjun Ma, Cihang Xie 等ICML 2026 · 被引用 2 次
