FastLR: Non-Autoregressive Lipreading Model with Integrate-and-Fire
Jinglin Liu, Yi Ren, Zhou Zhao, Chen Zhang, Baoxing Huai, Jing Yuan
摘要
Lipreading is an impressive technique and there has been a definite improvement of accuracy in recent years. However, existing methods for lipreading mainly build on autoregressive (AR) model, which generate target tokens one by one and suffer from high inference latency. To breakthrough this constraint, we propose FastLR, a non-autoregressive (NAR) lipreading model which generates all target tokens simultaneously. NAR lipreading is a challenging task that has many difficulties: 1) the discrepancy of sequence lengths between source and target makes it difficult to estimate the length of the output sequence; 2) the conditionally independent behavior of NAR generation lacks the correlation across time which leads to a poor approximation of target distribution; 3) the feature representation ability of encoder can be weak due to lack of effective alignment mechanism; and 4) the removal of AR language model exacerbates the inherent ambiguity problem of lipreading. Thus, in this paper, we introduce three methods to reduce the gap between FastLR and AR model: 1) to address challenges 1 and 2, we leverage integrate-and-fire (I&F) module to model the correspondence between source video frames and output text sequence. 2) To tackle challenge 3, we add an auxiliary connectionist temporal classification (CTC) decoder to the top of the encoder and optimize it with extra CTC loss. We also add an auxiliary autoregressive decoder to help the feature extraction of encoder. 3) To overcome challenge 4, we propose a novel Noisy Parallel Decoding (NPD) for I&F and bring Byte-Pair Encoding (BPE) into lipreading. Our experiments exhibit that FastLR achieves the speedup up to 10.97× comparing with state-of-the-art lipreading model with slight WER absolute increase of 1.5% and 5.5% on GRID and LRS2 lipreading datasets respectively, which demonstrates the effectiveness of our proposed method. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- SimulSLT: End-to-End Simultaneous Sign Language TranslationAoxiong Yin, Zhou Zhao, Jinglin Liu, Weike Jin 等ACM MM 2021 · 被引用 35 次
- SimulLR: Simultaneous Lip Reading Transducer with Attention-Guided Adaptive MemoryZhijie Lin, Zhou Zhao, Haoyuan Li, Jinglin Liu 等ACM MM 2021 · 被引用 13 次
- Parallel and High-Fidelity Text-to-Lip GenerationJinglin Liu, Zhiying Zhu, Yi Ren, Wencan Huang 等AAAI 2022 · 被引用 10 次
它引用的顶会 Paper5
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin 等ICLR 2021 · 被引用 513 次
- Hearing Lips: Improving Lip Reading by Distilling Speech RecognizersYa Zhao, Rui Xu, Xinchao Wang, Peng Hou 等AAAI 2020 · 被引用 106 次
- Non-Autoregressive Coarse-to-Fine Video CaptioningBang Yang, Yuexian Zou, Fenglin Liu, Can ZhangAAAI 2021 · 被引用 92 次
- Spatio-Temporal Fusion Based Convolutional Sequence Learning for Lip ReadingXingxuan Zhang, Feng Cheng, Shilin WangICCV 2019 · 被引用 87 次
- A Study of Non-autoregressive Model for Sequence GenerationYi Ren, Jinglin Liu, Xu Tan, Zhou Zhao 等ACL 2020 · 被引用 58 次
相关 Paper
- FastLTS: Non-Autoregressive End-to-End Unconstrained Lip-to-Speech SynthesisYongqi Wang, Zhou ZhaoACM MM 2022 · 被引用 9 次
- Flow-Based Unconstrained Lip to Speech GenerationJinzheng He, Zhou Zhao, Yi Ren, Jinglin Liu 等AAAI 2022 · 被引用 21 次
- VALLR: Visual ASR Language Model for Lip ReadingMarshall Thomas, Edward Fish, Richard BowdenICCV 2025 · 被引用 6 次
- An EM Approach to Non-autoregressive Conditional Sequence GenerationZhiqing Sun, Yiming YangICML 2020 · 被引用 43 次
- FastCorrect: Fast Error Correction with Edit Alignment for Automatic Speech RecognitionYichong Leng, Xu Tan, Linchen Zhu, Jin Xu 等NeurIPS 2021 · 被引用 84 次
