StyleLipSync: Style-based Personalized Lip-sync Video Generation
Taekyung Ki, Dongchan Min
Abstract
In this paper, we present StyleLipSync, a style-based personalized lip-sync video generative model that can generate identity-agnostic lip-synchronizing video from arbitrary audio. To generate a video of arbitrary identities, we leverage expressive lip prior from the semantically rich latent space of a pre-trained StyleGAN, where we can also design a video consistency with a linear transformation. In contrast to the previous lip-sync methods, we introduce pose-aware masking that dynamically locates the mask to improve the naturalness over frames by utilizing a 3D parametric mesh predictor frame by frame. Moreover, we propose a few-shot lip-sync adaptation method for an arbitrary person by introducing a sync regularizer that preserves lip-sync generalization while enhancing the person-specific visual information. Extensive experiments demonstrate that our model can generate accurate lip-sync videos even with the zero-shot setting and enhance characteristics of an unseen face using a few seconds of target video through the proposed adaptation method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext caab7302-9151-4800-a483-9f86d63d81c4Cited by top-tier papers6
- Lips Are Lying: Spotting the Temporal Inconsistency between Audio and Visual in Lip-Syncing DeepFakesWeifeng Liu, Tianyi She, Jiawei Liu, Boheng Li et al.NeurIPS 2024 · 57 citations
- Personalized Generation In Large Model Era: A SurveyYiyan Xu, Jinghao Zhang, Alireza Salemi, Xinting Hu et al.ACL 2025 · 45 citations
- Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural ConversationTaekyung Ki, Sangwon Jang, Jaehyeong Jo, Jaehong Yoon et al.CVPR 2026 · 18 citations
- FLOAT: Generative Motion Latent Flow Matching for Audio-Driven Talking PortraitTaekyung Ki, Dongchan Min, Gyeongsu ChaeICCV 2025 · 6 citations
- Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided DiffusionXingpei Ma, Jiaran Cai, Yuansheng Guan, Shenneng Huang et al.ICML 2025
Builds on18
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine et al.NeurIPS 2020 · 2,345 citations
- Alias-Free Generative Adversarial NetworksTero Karras, Miika Aittala, Samuli Laine, Erik Härkönen et al.NeurIPS 2021 · 2,126 citations
- Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space?Rameen Abdal, Yipeng Qin, Peter WonkaICCV 2019 · 1,195 citations
- GANSpace: Discovering Interpretable GAN ControlsErik Härkönen, Aaron Hertzmann, Jaakko Lehtinen, Sylvain ParisNeurIPS 2020 · 1,049 citations
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
Related papers
- StyleSync: High-Fidelity Generalized and Personalized Lip Sync in Style-Based GeneratorJiazhi Guan, Zhanwang Zhang, Hang Zhou, Tianshu Hu et al.CVPR 2023
- Enhancing Identity-Deformation Disentanglement in StyleGAN for One-Shot Face Video Re-EnactmentQing Chang, Yao-Xiang Ding, Kun ZhouAAAI 2025 · 3 citations
- Identity-Preserving Talking Face Generation with Landmark and Appearance PriorsWeizhi Zhong, Chaowei Fang, Yinqi Cai, Pengxu Wei et al.CVPR 2023
- StyleTalk: One-Shot Talking Head Generation with Controllable Speaking StylesYifeng Ma, Suzhen Wang, Zhipeng Hu, Changjie Fan et al.AAAI 2023 · 135 citations
- FlashLips: 100-FPS Mask-Free Latent Lip-Sync using Reconstruction Instead of Diffusion or GANsAndreas Zinonos, Michał Stypułkowski, Antoni Bigata Casademunt, Stavros Petridis et al.CVPR 2026
