CIEASR: Contextual Image-Enhanced Automatic Speech Recognition for Improved Homophone Discrimination
Ziyi Wang, Yiming Rong, Deyang Jiang, Haoran Wu, Shiyu Zhou, Bo Xu
Abstract
Automatic Speech Recognition (ASR) models pre-trained on large-scale speech datasets have achieved significant breakthroughs compared with traditional methods. However, mainstream pre-trained ASR models encounter challenges in distinguishing homophones, which have close or identical pronunciations. Previous studies have introduced visual auxiliary cues to address this challenge, yet the sophisticated use of lip movements falls short in correcting homophone errors. On the other hand, the fusion and utilization of scene images remain in an exploratory stage, with performance still inferior to the pre-trained speech model. In this paper, we introduce CIEASR (Contextual Image-Enhanced Automatic Speech Recognition), a novel multimodal speech recognition model that incorporates a new cue fusion method, using scene images as soft prompts to correct homophone errors. To mitigate data scarcity, we refine and expand the VSDial dataset for extensive experiments, illustrating that scene images contribute to the accurate recognition of entity nouns and personal pronouns. Our proposed CIEASR achieves state-of-the-art results on VSDial and Flickr8K, significantly reducing the Character Error Rate (CER) on VSDial from 3.61% to 0.92%.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers2
- Speech-Aware Long Context Pruning and Integration for Contextualized Automatic Speech RecognitionYiming Rong, Yixin Zhang, Ziyi Wang, Deyang Jiang et al.AAAI 2026
- Listening Like Humans: Semantics-Guided Noise-Robust Multimodal Speech RecognitionYan Fang, Jun Chen, Yian Yao, Shuxin Zhong et al.ACL 2026
Related papers
- VALLR: Visual ASR Language Model for Lip ReadingMarshall Thomas, Edward Fish, Richard BowdenICCV 2025 · 6 citations
- VHASR: A Multimodal Speech Recognition System With Vision HotwordsJiliang Hu, Zuchao Li, Ping Wang, Haojun Ai et al.EMNLP 2024
- AudioVSR: Enhancing Video Speech Recognition with Audio DataXiaoda Yang, Xize Cheng, Jiaqi Duan, Hongshun Qiu et al.EMNLP 2024 · 3 citations
- Extending Phrase Grounding with Pronouns in Visual DialoguesPanzhong Lu, Xin Zhang, Meishan Zhang, Min ZhangEMNLP 2022 · 5 citations
- Cuing Without Sharing: A Federated Cued Speech Recognition Framework via Mutual Knowledge DistillationYuxuan Zhang, Lei Liu, Li LiuACM MM 2023 · 8 citations
