Bridging the Gap: Using Deep Acoustic Representations to Learn Grounded Language from Percepts and Raw Speech
Gaoussou Youssouf Kebe, Luke E. Richards, Edward Raff, Francis Ferraro, Cynthia Matuszek
摘要
Learning to understand grounded language, which connects natural language to percepts, is a critical research area. Prior work in grounded language acquisition has focused primarily on textual inputs. In this work, we demonstrate the feasibility of performing grounded language acquisition on paired visual percepts and raw speech inputs. This will allow human-robot interactions in which language about novel tasks and environments is learned from end-users, reducing dependence on textual inputs and potentially mitigating the effects of demographic bias found in widely available speech recognition systems. We leverage recent work in self-supervised speech representation models and show that learned representations of speech can make language grounding systems more inclusive towards specific groups while maintaining or even increasing general performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- vq-wav2vec: Self-Supervised Learning of Discrete Speech RepresentationsAlexei Baevski, Steffen Schneider, Michael AuliICLR 2020 · 被引用 730 次
- It's Morphin' Time! Combating Linguistic Discrimination with Inflectional PerturbationsSamson Tan, Shafiq R. Joty, Min-Yen Kan, Richard SocherACL 2020 · 被引用 88 次
相关 Paper
- World-to-Words: Grounded Open Vocabulary Acquisition through Fast Mapping in Vision-Language ModelsZiqiao Ma, Jiayi Pan, Joyce ChaiACL 2023 · 被引用 4 次
- Learning to Ground Multi-Agent Communication with AutoencodersToru Lin, Jacob Huh, Christopher Stauffer, Ser-Nam Lim 等NeurIPS 2021 · 被引用 75 次
- Grounding Spatio-Temporal Language with TransformersTristan Karch, Laetitia Teodorescu, Katja Hofmann, Clément Moulin-Frier 等NeurIPS 2021 · 被引用 11 次
- Walk in Others' Shoes with a Single Glance: Human-Centric Visual Grounding with Top-View Perspective TransformationYuqi Bu, Xin Wu, Zirui Zhao, Yi Cai 等ACL 2025
- What Do Language Models Hear? Probing for Auditory Representations in Language ModelsJerry Ngo, Yoon KimACL 2024
