Parallel and High-Fidelity Text-to-Lip Generation
Jinglin Liu, Zhiying Zhu, Yi Ren, Wencan Huang, Baoxing Huai, Nicholas Jing Yuan, Zhou Zhao
摘要
As a key component of talking face generation, lip movements generation determines the naturalness and coherence of the generated talking face video. Prior literature mainly focuses on speech-to-lip generation while there is a paucity in text-to-lip (T2L) generation. T2L is a challenging task and existing end-to-end works depend on the attention mechanism and autoregressive (AR) decoding manner. However, the AR decoding manner generates current lip frame conditioned on frames generated previously, which inherently hinders the inference speed, and also has a detrimental effect on the quality of generated lip frames due to error propagation. This encourages the research of parallel T2L generation. In this work, we propose a parallel decoding model for fast and high-fidelity text-to-lip generation (ParaLip). Specifically, we predict the duration of the encoded linguistic features and model the target lip frames conditioned on the encoded linguistic features with their duration in a non-autoregressive manner. Furthermore, we incorporate the structural similarity index loss and adversarial learning to improve perceptual quality of generated lip frames and alleviate the blurry prediction problem. Extensive experiments conducted on GRID and TCD-TIMIT datasets demonstrate the superiority of proposed methods. Video samples are available via https://paralip.github.io/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- DINet: Deformation Inpainting Network for Realistic Face Visually Dubbing on High Resolution VideoZhimeng Zhang, Zhipeng Hu, Wenjin Deng, Changjie Fan 等AAAI 2023 · 被引用 106 次
- AV-Flow: Transforming Text to Audio-Visual Human-Like InteractionsAggelina Chatziagapi, Louis-Philippe Morency, Hongyu Gong, Michael Zollhöfer 等ICCV 2025 · 被引用 3 次
- Faces that Speak: Jointly Synthesising Talking Face and Speech from TextYoungjoon Jang, Ji-Hoon Kim, Junseok Ahn, Doyeop Kwak 等CVPR 2024
它引用的顶会 Paper5
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 被引用 869 次
- Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment SearchJaehyeon Kim, Sungwon Kim, Jungil Kong, Sungroh YoonNeurIPS 2020 · 被引用 663 次
- DualLip: A System for Joint Lip Reading and GenerationWeicong Chen, Xu Tan, Yingce Xia, Tao Qin 等ACM MM 2020 · 被引用 25 次
- FastLR: Non-Autoregressive Lipreading Model with Integrate-and-FireJinglin Liu, Yi Ren, Zhou Zhao, Chen Zhang 等ACM MM 2020 · 被引用 13 次
- SimulLR: Simultaneous Lip Reading Transducer with Attention-Guided Adaptive MemoryZhijie Lin, Zhou Zhao, Haoyuan Li, Jinglin Liu 等ACM MM 2021 · 被引用 13 次
相关 Paper
- Seeing What You Said: Talking Face Generation Guided by a Lip Reading ExpertJiadong Wang, Xinyuan Qian, Malu Zhang, Robby T. Tan 等CVPR 2023
- SyncTalklip: Highly Synchronized Lip-Readable Speaker Generation with Multi-Task LearningXiaoda Yang, Xize Cheng, Dongjie Fu, Minghui Fang 等ACM MM 2024 · 被引用 4 次
- SyncTalkFace: Talking Face Generation with Precise Lip-Syncing via Audio-Lip MemorySe Jin Park, Minsu Kim, Joanna Hong, Jeongsoo Choi 等AAAI 2022 · 被引用 110 次
- Speech2Lip: High-fidelity Speech to Lip Generation by Learning from a Short VideoXiuzhe Wu, Pengfei Hu, Yang Wu, Xiaoyang Lyu 等ICCV 2023 · 被引用 18 次
- FastLTS: Non-Autoregressive End-to-End Unconstrained Lip-to-Speech SynthesisYongqi Wang, Zhou ZhaoACM MM 2022 · 被引用 9 次
