Parallel and High-Fidelity Text-to-Lip Generation
Jinglin Liu, Zhiying Zhu, Yi Ren, Wencan Huang, Baoxing Huai, Nicholas Jing Yuan, Zhou Zhao
Abstract
As a key component of talking face generation, lip movements generation determines the naturalness and coherence of the generated talking face video. Prior literature mainly focuses on speech-to-lip generation while there is a paucity in text-to-lip (T2L) generation. T2L is a challenging task and existing end-to-end works depend on the attention mechanism and autoregressive (AR) decoding manner. However, the AR decoding manner generates current lip frame conditioned on frames generated previously, which inherently hinders the inference speed, and also has a detrimental effect on the quality of generated lip frames due to error propagation. This encourages the research of parallel T2L generation. In this work, we propose a parallel decoding model for fast and high-fidelity text-to-lip generation (ParaLip). Specifically, we predict the duration of the encoded linguistic features and model the target lip frames conditioned on the encoded linguistic features with their duration in a non-autoregressive manner. Furthermore, we incorporate the structural similarity index loss and adversarial learning to improve perceptual quality of generated lip frames and alleviate the blurry prediction problem. Extensive experiments conducted on GRID and TCD-TIMIT datasets demonstrate the superiority of proposed methods. Video samples are available via https://paralip.github.io/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4cb8516a-3388-4c72-8776-d7e49c4731c8Cited by top-tier papers3
- DINet: Deformation Inpainting Network for Realistic Face Visually Dubbing on High Resolution VideoZhimeng Zhang, Zhipeng Hu, Wenjin Deng, Changjie Fan et al.AAAI 2023 · 106 citations
- AV-Flow: Transforming Text to Audio-Visual Human-Like InteractionsAggelina Chatziagapi, Louis-Philippe Morency, Hongyu Gong, Michael Zollhöfer et al.ICCV 2025 · 3 citations
- Faces that Speak: Jointly Synthesising Talking Face and Speech from TextYoungjoon Jang, Ji-Hoon Kim, Junseok Ahn, Doyeop Kwak et al.CVPR 2024
Builds on5
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment SearchJaehyeon Kim, Sungwon Kim, Jungil Kong, Sungroh YoonNeurIPS 2020 · 663 citations
- DualLip: A System for Joint Lip Reading and GenerationWeicong Chen, Xu Tan, Yingce Xia, Tao Qin et al.ACM MM 2020 · 25 citations
- FastLR: Non-Autoregressive Lipreading Model with Integrate-and-FireJinglin Liu, Yi Ren, Zhou Zhao, Chen Zhang et al.ACM MM 2020 · 13 citations
- SimulLR: Simultaneous Lip Reading Transducer with Attention-Guided Adaptive MemoryZhijie Lin, Zhou Zhao, Haoyuan Li, Jinglin Liu et al.ACM MM 2021 · 13 citations
Related papers
- Seeing What You Said: Talking Face Generation Guided by a Lip Reading ExpertJiadong Wang, Xinyuan Qian, Malu Zhang, Robby T. Tan et al.CVPR 2023
- SyncTalklip: Highly Synchronized Lip-Readable Speaker Generation with Multi-Task LearningXiaoda Yang, Xize Cheng, Dongjie Fu, Minghui Fang et al.ACM MM 2024 · 4 citations
- SyncTalkFace: Talking Face Generation with Precise Lip-Syncing via Audio-Lip MemorySe Jin Park, Minsu Kim, Joanna Hong, Jeongsoo Choi et al.AAAI 2022 · 110 citations
- Speech2Lip: High-fidelity Speech to Lip Generation by Learning from a Short VideoXiuzhe Wu, Pengfei Hu, Yang Wu, Xiaoyang Lyu et al.ICCV 2023 · 18 citations
- FastLTS: Non-Autoregressive End-to-End Unconstrained Lip-to-Speech SynthesisYongqi Wang, Zhou ZhaoACM MM 2022 · 9 citations
