Occlusion-Insensitive Talking Head Video Generation via Facelet Compensation
Yuhui Deng, Yuqin Lu, Yangyang Xu, Yongwei Nie, Shengfeng He
Abstract
Talking head video generation involves animating a still face image using facial motion cues derived from a driving video to replicate target poses and expressions. Traditional methods often rely on the assumption that the relative positions of facial keypoints remain unchanged. However, this assumption fails when keypoints are occluded or when the head is in a profile pose, leading to inconsistencies in identity and blurring in certain facial regions. In this paper, we introduce Occlusion-Insensitive Talking Head Video Generation, a novel approach that eliminates the reliance on spatial correlation of keypoints and instead leverages semantic correlation. Our method transforms facial features into a facelet semantic bank, where each facelet token represents a specific facial semantic. This bank is devoid of spatial information, allowing it to compensate for any invisible or occluded face regions during motion warping. The facelet compensation module then populates the facelet tokens within the initially warped features by learning a correlation matrix between facial semantics and the facelet bank. This approach enables precise compensation for occlusions and pose changes, enhancing the fidelity of the generated videos. Extensive experiments demonstrate that our method achieves state-of-the-art results, preserving source identity, maintaining fine-grained facial details, and capturing nuanced facial expressions with remarkable accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f1aaee2c-b915-4f87-a523-23658f083098Builds on20
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- Few-Shot Adversarial Learning of Realistic Neural Talking Head ModelsEgor Zakharov, Aliaksandra Shysheya, Egor Burkov, Victor S. LempitskyICCV 2019 · 687 citations
- Learning an animatable detailed 3D face model from in-the-wild imagesYao Feng, Haiwen Feng, Michael J. Black, Timo BolkartSIGGRAPH 2021 · 662 citations
- AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisYudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu et al.ICCV 2021 · 510 citations
- Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion ModelsHila Chefer, Yuval Alaluf, Yael Vinker, Lior Wolf et al.SIGGRAPH 2023 · 438 citations
Related papers
- Implicit Identity Representation Conditioned Memory Compensation Network for Talking Head Video GenerationFa-Ting Hong, Dan XuICCV 2023 · 75 citations
- Synergizing Motion and Appearance: Multi-Scale Compensatory Codebooks for Talking Head Video GenerationShuling Zhao, Fa-Ting Hong, Xiaoshui Huang, Dan XuCVPR 2025
- Depth-Aware Generative Adversarial Network for Talking Head Video GenerationFa-Ting Hong, Longhao Zhang, Li Shen, Dan XuCVPR 2022 · 168 citations
- Write-a-speaker: Text-based Emotional and Rhythmic Talking-head GenerationLincheng Li, Suzhen Wang, Zhimeng Zhang, Yu Ding et al.AAAI 2021 · 88 citations
- Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual RepresentationHang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy et al.CVPR 2021
