A Lip Sync Expert Is All You Need for Speech to Lip Generation In the Wild
K. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. Jawahar
Abstract
In this work, we investigate the problem of lip-syncing a talking face video of an arbitrary identity to match a target speech segment. Current works excel at producing accurate lip movements on a static image or videos of specific people seen during the training phase. However, they fail to accurately morph the lip movements of arbitrary identities in dynamic, unconstrained talking face videos, resulting in significant parts of the video being out-of-sync with the new audio. We identify key reasons pertaining to this and hence resolve them by learning from a powerful lip-sync discriminator. Next, we propose new, rigorous evaluation benchmarks and metrics to accurately measure lip synchronization in unconstrained videos. Extensive quantitative evaluations on our challenging benchmarks show that the lip-sync accuracy of the videos generated by our Wav2Lip model is almost as good as real synced videos. We provide a demo video clearly showing the substantial impact of our Wav2Lip model, and also publicly release the code, models, and evaluation benchmarks on our website.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e3635e44-c77d-42fa-ae78-1af864d38120Cited by top-tier papers203
- MeshTalk: 3D Face Animation from Speech using Cross-Modality DisentanglementAlexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre et al.ICCV 2021 · 272 citations
- VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeSicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang et al.NeurIPS 2024 · 253 citations
- FaceFormer: Speech-Driven 3D Facial Animation with TransformersYingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang et al.CVPR 2022 · 218 citations
- EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark ConditionsZhiyuan Chen, Jiajiong Cao, Zhiquan Chen, Yuming Li et al.AAAI 2025 · 197 citations
- KoDF: A Large-scale Korean DeepFake Detection DatasetPatrick Kwon, Jaeseong You, Gyuhyeon Nam, Sungwoo Park et al.ICCV 2021 · 154 citations
Related papers
- Seeing What You Said: Talking Face Generation Guided by a Lip Reading ExpertJiadong Wang, Xinyuan Qian, Malu Zhang, Robby T. Tan et al.CVPR 2023
- Learning Individual Speaking Styles for Accurate Lip to Speech SynthesisK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharCVPR 2020
- Lip-to-Speech Synthesis for Arbitrary Speakers in the WildSindhu B. Hegde, K. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri et al.ACM MM 2022 · 15 citations
- SIDGAN: High-Resolution Dubbed Video Generation via Shift-Invariant LearningUrwa Muaz, Wondong Jang, Rohun Tripathi, Santhosh Mani et al.ICCV 2023 · 8 citations
- Speech2Lip: High-fidelity Speech to Lip Generation by Learning from a Short VideoXiuzhe Wu, Pengfei Hu, Yang Wu, Xiaoyang Lyu et al.ICCV 2023 · 18 citations
