Bridging the Gap: Using Deep Acoustic Representations to Learn Grounded Language from Percepts and Raw Speech
Gaoussou Youssouf Kebe, Luke E. Richards, Edward Raff, Francis Ferraro, Cynthia Matuszek
Abstract
Learning to understand grounded language, which connects natural language to percepts, is a critical research area. Prior work in grounded language acquisition has focused primarily on textual inputs. In this work, we demonstrate the feasibility of performing grounded language acquisition on paired visual percepts and raw speech inputs. This will allow human-robot interactions in which language about novel tasks and environments is learned from end-users, reducing dependence on textual inputs and potentially mitigating the effects of demographic bias found in widely available speech recognition systems. We leverage recent work in self-supervised speech representation models and show that learned representations of speech can make language grounding systems more inclusive towards specific groups while maintaining or even increasing general performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c8dc8b8d-9c98-4e61-b95c-d6557772efa4Builds on3
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- vq-wav2vec: Self-Supervised Learning of Discrete Speech RepresentationsAlexei Baevski, Steffen Schneider, Michael AuliICLR 2020 · 730 citations
- It's Morphin' Time! Combating Linguistic Discrimination with Inflectional PerturbationsSamson Tan, Shafiq R. Joty, Min-Yen Kan, Richard SocherACL 2020 · 88 citations
Related papers
- World-to-Words: Grounded Open Vocabulary Acquisition through Fast Mapping in Vision-Language ModelsZiqiao Ma, Jiayi Pan, Joyce ChaiACL 2023 · 4 citations
- Learning to Ground Multi-Agent Communication with AutoencodersToru Lin, Jacob Huh, Christopher Stauffer, Ser-Nam Lim et al.NeurIPS 2021 · 75 citations
- Grounding Spatio-Temporal Language with TransformersTristan Karch, Laetitia Teodorescu, Katja Hofmann, Clément Moulin-Frier et al.NeurIPS 2021 · 11 citations
- Walk in Others' Shoes with a Single Glance: Human-Centric Visual Grounding with Top-View Perspective TransformationYuqi Bu, Xin Wu, Zirui Zhao, Yi Cai et al.ACL 2025
- What Do Language Models Hear? Probing for Auditory Representations in Language ModelsJerry Ngo, Yoon KimACL 2024
