Expressive Talking Head Generation with Granular Audio-Visual Control
Borong Liang, Yan Pan, Zhizhi Guo, Hang Zhou, Zhibin Hong, Xiaoguang Han, Junyu Han, Jingtuo Liu, Errui Ding, Jingdong Wang
Abstract
Generating expressive talking heads is essential for creating virtual humans. However, existing one- or few-shot methods focus on lip-sync and head motion, ignoring the emotional expressions that make talking faces realistic. In this paper, we propose the Granularly Controlled Audio-Visual Talking Heads (GC-AVT), which controls lip movements, head poses, and facial expressions of a talking head in a granular manner. Our insight is to decouple the audio-visual driving sources through prior-based pre-processing designs. Detailedly, we disassemble the driving image into three complementary parts including: 1) a cropped mouth that facilitates lip-sync; 2) a masked head that implicitly learns pose; and 3) the upper face which works corporately and complementarily with a time-shifted mouth to contribute the expression. Interestingly, the encoded features from the three sources are integrally balanced through reconstruction training. Extensive experiments show that our method generates expressive faces with not only synced mouth shapes, controllable poses, but precisely animated emotional expressions as well.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 25f425c9-7fd6-4bab-bc32-d8574befcafdCited by top-tier papers41
- VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeSicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang et al.NeurIPS 2024 · 253 citations
- StyleTalk: One-Shot Talking Head Generation with Controllable Speaking StylesYifeng Ma, Suzhen Wang, Zhipeng Hu, Changjie Fan et al.AAAI 2023 · 135 citations
- Efficient Emotional Adaptation for Audio-Driven Talking-Head GenerationYuan Gan, Zongxin Yang, Xihang Yue, Lingyun Sun et al.ICCV 2023 · 111 citations
- DINet: Deformation Inpainting Network for Realistic Face Visually Dubbing on High Resolution VideoZhimeng Zhang, Zhipeng Hu, Wenjin Deng, Changjie Fan et al.AAAI 2023 · 106 citations
- Implicit Identity Representation Conditioned Memory Compensation Network for Talking Head Video GenerationFa-Ting Hong, Dan XuICCV 2023 · 75 citations
Builds on14
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- Few-Shot Adversarial Learning of Realistic Neural Talking Head ModelsEgor Zakharov, Aliaksandra Shysheya, Egor Burkov, Victor S. LempitskyICCV 2019 · 687 citations
- HeadGAN: One-shot Neural Head Synthesis and EditingMichail Christos Doukas, Stefanos Zafeiriou, Viktoriia SharmanskaICCV 2021 · 164 citations
- Learning Hierarchical Cross-Modal Association for Co-Speech Gesture GenerationXian Liu, Qianyi Wu, Hang Zhou, Yinghao Xu et al.CVPR 2022 · 118 citations
- Vision-Infused Deep Audio InpaintingHang Zhou, Ziwei Liu, Xudong Xu, Ping Luo et al.ICCV 2019 · 92 citations
Related papers
- Progressive Disentangled Representation Learning for Fine-Grained Controllable Talking Head SynthesisDuomin Wang, Yu Deng, Zixin Yin, Heung-Yeung Shum et al.CVPR 2023
- MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled EmbeddingChang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja et al.ACM MM 2025 · 3 citations
- FG-EmoTalk: Talking Head Video Generation with Fine-Grained Controllable Facial ExpressionsZhaoxu Sun, Yuze Xuan, Fang Liu, Yang XiangAAAI 2024 · 13 citations
- AUHead: Realistic Emotional Talking Head Generation via Action Units ControlJiayi Lyu, Leigang Qu, Wenjing Zhang, Hanyu Jiang et al.ICLR 2026 · 2 citations
- PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face GenerationBaiqin Wang, Xiangyu Zhu, Fan Shen, Hao Xu et al.CVPR 2026 · 8 citations
