That's What I Said: Fully-Controllable Talking Face Generation
Youngjoon Jang, Kyeongha Rho, Jong-Bin Woo, Hyeongkeun Lee, Jihwan Park, Youshin Lim, Byeong-Yeol Kim, Joon Son Chung
Abstract
The goal of this paper is to synthesise talking faces with controllable facial motions. To achieve this goal, we propose two key ideas. The first is to establish a canonical space where every face has the same motion patterns but different identities. The second is to navigate a multimodal motion space that only represents motion-related features while eliminating identity information. To disentangle identity and motion, we introduce an orthogonality constraint between the two different latent spaces. From this, our method can generate natural-looking talking faces with fully controllable facial attributes and accurate lip synchronisation. Extensive experiments demonstrate that our method achieves state-of-the-art results in terms of both visual quality and lip-sync score. To the best of our knowledge, we are the first to develop a talking face generation framework that can accurately manifest full target facial motions including lip, head pose, and eye movements in the generated video without any additional supervision beyond RGB video with audio.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on19
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- Unsupervised Discovery of Interpretable Directions in the GAN Latent SpaceAndrey Voynov, Artem BabenkoICML 2020 · 459 citations
- On the "steerability" of generative adversarial networksAli Jahanian, Lucy Chai, Phillip IsolaICLR 2020 · 421 citations
- GANalyze: Toward Visual Definitions of Cognitive Image PropertiesLore Goetschalckx, Alex Andonian, Aude Oliva, Phillip IsolaICCV 2019 · 345 citations
- HyperStyle: StyleGAN Inversion with HyperNetworks for Real Image EditingYuval Alaluf, Omer Tov, Ron Mokady, Rinon Gal et al.CVPR 2022 · 250 citations
Related papers
- Progressive Disentangled Representation Learning for Fine-Grained Controllable Talking Head SynthesisDuomin Wang, Yu Deng, Zixin Yin, Heung-Yeung Shum et al.CVPR 2023
- DGTalker: Disentangled Generative Latent Space Learning for Audio-Driven Gaussian Talking HeadsXiaoxi Liang, Yanbo Fan, Qiya Yang, Xuan Wang et al.ICCV 2025 · 2 citations
- Talking Face Generation with Expression-Tailored Generative Adversarial NetworkDan Zeng, Han Liu, Hui Lin, Shiming GeACM MM 2020 · 30 citations
- FACIAL: Synthesizing Dynamic Talking Face with Implicit Attribute LearningChenxu Zhang, Yifan Zhao, Yifei Huang, Ming Zeng et al.ICCV 2021 · 149 citations
- PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face GenerationBaiqin Wang, Xiangyu Zhu, Fan Shen, Hao Xu et al.CVPR 2026 · 8 citations
