Cueing Without Gapping: Cuer-Independent Cued Speech Recognition Powered by Cross-Cuer Invariant Modeling
Fengji Ma, Chenxing Li, Li Liu
Abstract
Automatic Cued Speech Recognition (ACSR) is a vital communication system designed to enhance spoken language accessibility for the hearing-impaired by combining lip movements and hand gestures to encode phonemes. Despite its effectiveness, current ACSR methods face significant challenges, including poor generalization to unseen cuers 1 due to the limited scale of CS datasets, which restricts the ability of existing visual encoder to capture cuer-invariant CS visual features. Additionally, previous approaches relying on Connectionist Temporal Classification (CTC) decoding fail to incorporate prior linguistic sequence knowledge, further limiting their performance. To address these issues, we propose a novel Two Auxiliary Modalities guided Cross-cuer Invariant Adaptation method (TACIA), introducing pose and text modalities to help extract cuer-invariant motion and semantic features, thereby improving generalization. In addition, we introduce a Visual-guided Cued Token Prediction (VG-NTP) method, inspired by large language models. This method replaces CTC decoding by incorporating language modeling, leveraging rich linguistic knowledge, including semantics, to address the suboptimal issues present in the CTC decoding process. Extensive experiments demonstrate the superiority of our approach to the state-of-the-art (SOTA) on Chinese and British CS datasets, significantly advancing the accuracy and quality of ACSR systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4f4b5629-2e05-4790-9d6d-d6da4b22120aBuilds on5
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Gloss-free Sign Language Translation: Improving from Visual-Language PretrainingBenjia Zhou, Zhigang Chen, Albert Clapés, Jun Wan et al.ICCV 2023 · 123 citations
- Sign Language Transformers: Joint End-to-End Sign Language Recognition and TranslationNecati Cihan Camgöz, Oscar Koller, Simon Hadfield, Richard BowdenCVPR 2020
- Bringing Inputs to Shared Domains for 3D Interacting Hands Recovery in the WildGyeongsik MoonCVPR 2023
- Uni-Sign: Toward Unified Sign Language Understanding at ScaleZecheng Li, Wengang Zhou, Weichao Zhao, Kepeng Wu et al.ICLR 2025
Related papers
- Cued-Agent: A Collaborative Multi-Agent System for Automatic Cued Speech RecognitionGuanjie Huang, Danny H. K. Tsang, Shan Yang, Guangzhi Lei et al.ACM MM 2025
- Cuing Without Sharing: A Federated Cued Speech Recognition Framework via Mutual Knowledge DistillationYuxuan Zhang, Lei Liu, Li LiuACM MM 2023 · 8 citations
- UniCUE: Unified Recognition and Generation Framework for Chinese Cued Speech Video-to-Speech GenerationJinting Wang, Shan Yang, Chenxing Li, Dong Yu et al.AAAI 2026
- Bridging Sign and Spoken Languages: Pseudo Gloss Generation for Sign Language TranslationJianyuan Guo, Peike Li, Trevor CohnNeurIPS 2025 · 17 citations
- VALLR: Visual ASR Language Model for Lip ReadingMarshall Thomas, Edward Fish, Richard BowdenICCV 2025 · 6 citations
