FastLR: Non-Autoregressive Lipreading Model with Integrate-and-Fire
Jinglin Liu, Yi Ren, Zhou Zhao, Chen Zhang, Baoxing Huai, Jing Yuan
Abstract
Lipreading is an impressive technique and there has been a definite improvement of accuracy in recent years. However, existing methods for lipreading mainly build on autoregressive (AR) model, which generate target tokens one by one and suffer from high inference latency. To breakthrough this constraint, we propose FastLR, a non-autoregressive (NAR) lipreading model which generates all target tokens simultaneously. NAR lipreading is a challenging task that has many difficulties: 1) the discrepancy of sequence lengths between source and target makes it difficult to estimate the length of the output sequence; 2) the conditionally independent behavior of NAR generation lacks the correlation across time which leads to a poor approximation of target distribution; 3) the feature representation ability of encoder can be weak due to lack of effective alignment mechanism; and 4) the removal of AR language model exacerbates the inherent ambiguity problem of lipreading. Thus, in this paper, we introduce three methods to reduce the gap between FastLR and AR model: 1) to address challenges 1 and 2, we leverage integrate-and-fire (I&F) module to model the correspondence between source video frames and output text sequence. 2) To tackle challenge 3, we add an auxiliary connectionist temporal classification (CTC) decoder to the top of the encoder and optimize it with extra CTC loss. We also add an auxiliary autoregressive decoder to help the feature extraction of encoder. 3) To overcome challenge 4, we propose a novel Noisy Parallel Decoding (NPD) for I&F and bring Byte-Pair Encoding (BPE) into lipreading. Our experiments exhibit that FastLR achieves the speedup up to 10.97× comparing with state-of-the-art lipreading model with slight WER absolute increase of 1.5% and 5.5% on GRID and LRS2 lipreading datasets respectively, which demonstrates the effectiveness of our proposed method. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0f367e69-30eb-4cc9-98bf-322452ab2718Cited by top-tier papers3
- SimulSLT: End-to-End Simultaneous Sign Language TranslationAoxiong Yin, Zhou Zhao, Jinglin Liu, Weike Jin et al.ACM MM 2021 · 35 citations
- SimulLR: Simultaneous Lip Reading Transducer with Attention-Guided Adaptive MemoryZhijie Lin, Zhou Zhao, Haoyuan Li, Jinglin Liu et al.ACM MM 2021 · 13 citations
- Parallel and High-Fidelity Text-to-Lip GenerationJinglin Liu, Zhiying Zhu, Yi Ren, Wencan Huang et al.AAAI 2022 · 10 citations
Builds on5
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin et al.ICLR 2021 · 513 citations
- Hearing Lips: Improving Lip Reading by Distilling Speech RecognizersYa Zhao, Rui Xu, Xinchao Wang, Peng Hou et al.AAAI 2020 · 106 citations
- Non-Autoregressive Coarse-to-Fine Video CaptioningBang Yang, Yuexian Zou, Fenglin Liu, Can ZhangAAAI 2021 · 92 citations
- Spatio-Temporal Fusion Based Convolutional Sequence Learning for Lip ReadingXingxuan Zhang, Feng Cheng, Shilin WangICCV 2019 · 87 citations
- A Study of Non-autoregressive Model for Sequence GenerationYi Ren, Jinglin Liu, Xu Tan, Zhou Zhao et al.ACL 2020 · 58 citations
Related papers
- FastLTS: Non-Autoregressive End-to-End Unconstrained Lip-to-Speech SynthesisYongqi Wang, Zhou ZhaoACM MM 2022 · 9 citations
- Flow-Based Unconstrained Lip to Speech GenerationJinzheng He, Zhou Zhao, Yi Ren, Jinglin Liu et al.AAAI 2022 · 21 citations
- VALLR: Visual ASR Language Model for Lip ReadingMarshall Thomas, Edward Fish, Richard BowdenICCV 2025 · 6 citations
- An EM Approach to Non-autoregressive Conditional Sequence GenerationZhiqing Sun, Yiming YangICML 2020 · 43 citations
- FastCorrect: Fast Error Correction with Edit Alignment for Automatic Speech RecognitionYichong Leng, Xu Tan, Linchen Zhu, Jin Xu et al.NeurIPS 2021 · 84 citations
