Progressive Disentangled Representation Learning for Fine-Grained Controllable Talking Head Synthesis
Duomin Wang, Yu Deng, Zixin Yin, Heung-Yeung Shum, Baoyuan Wang
2023Year
44Top-tier citations
Abstract
Figure 1. Our method takes an appearance reference as input and generates its talking head with disentangled control over lip motion, head pose, eye gaze&blink, and emotional expression, where the driving signal of lip motion comes from speech audio, and all other motions are controlled by different videos. As shown, it well disentangles all motion factors and achieves precise control over individual motion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers44
- VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeSicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang et al.NeurIPS 2024 · 253 citations
- HumanTOMATO: Text-aligned Whole-body Motion GenerationShunlin Lu, Ling-Hao Chen, Ailing Zeng, Jing Lin et al.ICML 2024 · 124 citations
- Generalizable and Animatable Gaussian Head AvatarXuangeng Chu, Tatsuya HaradaNeurIPS 2024 · 115 citations
- GAIA: Zero-shot Talking Avatar GenerationTianyu He, Junliang Guo, Runyi Yu, Yuchi Wang et al.ICLR 2024 · 51 citations
- SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human GenerationYouliang Zhang, Zhaoyang Li, Duomin Wang, jiahe zhang et al.ICLR 2026 · 30 citations
Builds on25
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- PIRenderer: Controllable Portrait Image Generation via Semantic Neural RenderingYurui Ren, Ge Li, Yuanqi Chen, Thomas H. Li et al.ICCV 2021 · 284 citations
- Latent Image Animator: Learning to Animate Images via Latent Space NavigationYaohui Wang, Di Yang, François Brémond, Antitza DantchevaICLR 2022 · 219 citations
- EMOCA: Emotion Driven Monocular Face Capture and AnimationRadek Danecek, Michael J. Black, Timo BolkartCVPR 2022 · 180 citations
- Depth-Aware Generative Adversarial Network for Talking Head Video GenerationFa-Ting Hong, Longhao Zhang, Li Shen, Dan XuCVPR 2022 · 168 citations
Related papers
- That's What I Said: Fully-Controllable Talking Face GenerationYoungjoon Jang, Kyeongha Rho, Jong-Bin Woo, Hyeongkeun Lee et al.ACM MM 2023 · 7 citations
- Expressive Talking Head Generation with Granular Audio-Visual ControlBorong Liang, Yan Pan, Zhizhi Guo, Hang Zhou et al.CVPR 2022 · 114 citations
- GoHD: Gaze-oriented and Highly Disentangled Portrait Animation with Rhythmic Poses and Realistic ExpressionsZiqi Zhou, Weize Quan, Hailin Shi, Wei Li et al.AAAI 2025 · 1 citation
- Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual RepresentationHang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy et al.CVPR 2021
- MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled EmbeddingChang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja et al.ACM MM 2025 · 3 citations
