Can Visual Context Improve Automatic Speech Recognition for an Embodied Agent?
Pradip Pramanick, Chayan Sarkar
摘要
The usage of automatic speech recognition (ASR) systems are becoming omnipresent ranging from personal assistant to chatbots, home, and industrial automation systems, etc. Modern robots are also equipped with ASR capabilities for interacting with humans as speech is the most natural interaction modality. However, ASR in robots faces additional challenges as compared to a personal assistant. Being an embodied agent, a robot must recognize the physical entities around it and therefore reliably recognize the speech containing the description of such entities. However, current ASR systems are often unable to do so due to limitations in ASR training, such as generic datasets and open-vocabulary modeling. Also, adverse conditions during inference, such as noise, accented, and far-field speech makes the transcription inaccurate. In this work, we present a method to incorporate a robot's visual information into an ASR system and improve the recognition of a spoken utterance containing a visible entity. Specifically, we propose a new decoder biasing technique to incorporate the visual context while ensuring the ASR output does not degrade for incorrect context. We achieve a 59% relative reduction in WER from an unmodified ASR system.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- VAPO: End-to-end Slide-Enhanced Speech Recognition with Omni-modal Large Language ModelsRui Hu, Delai Qiu, Yining Wang, Shengping Liu 等ACL 2026
- Leveraging Modality-Specific Representations for Audio-Visual Speech Recognition via Reinforcement LearningChen Chen, Yuchen Hu, Qiang Zhang, Heqing Zou 等AAAI 2023 · 被引用 35 次
- Listening Like Humans: Semantics-Guided Noise-Robust Multimodal Speech RecognitionYan Fang, Jun Chen, Yian Yao, Shuxin Zhong 等ACL 2026
- Do Slides Help? Multi-modal Context for Automatic Transcription of Conference TalksSupriti Sinhamahapatra, Jan NiehuesEMNLP 2025
- Enhancing Audiovisual Speech Recognition Through Bifocal Preference OptimizationYihan Wu, Yichen Lu, Yifan Peng, Xihua Wang 等AAAI 2025 · 被引用 1 次
