Implicit Identity Representation Conditioned Memory Compensation Network for Talking Head Video Generation
Fa-Ting Hong, Dan Xu
Abstract
Talking head video generation aims to animate a human face in a still image with dynamic poses and expressions using motion information derived from a target-driving video, while maintaining the person’s identity in the source image. However, dramatic and complex motions in the driving video cause ambiguous generation, because the still source image cannot provide sufficient appearance information for occluded regions or delicate expression variations, which produces severe artifacts and significantly degrades the generation quality. To tackle this problem, we propose to learn a global facial representation space, and design a novel implicit identity representation conditioned memory compensation network, coined as MCNet, for high-fidelity talking head generation. Specifically, we devise a network module to learn a unified spatial facial meta-memory bank from all training samples, which can provide rich facial structure and appearance priors to compensate warped source facial features for the generation. Furthermore, we propose an effective query mechanism based on implicit identity representations learned from the discrete keypoints of the source image. It can greatly facilitate the retrieval of more correlated information from the memory bank for the compensation. Extensive experiments demonstrate that MCNet can learn representative and complementary facial memory, and can clearly outperform previous state-of-the-art talking head generation methods on VoxCeleb1 and CelebV datasets. Please check our Project.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 70134709-52fd-4559-afc4-1fcc2d589c2fCited by top-tier papers20
- Real3D-Portrait: One-shot Realistic 3D Talking Portrait SynthesisZhenhui Ye, Tianyun Zhong, Yi Ren, Jiaqi Yang et al.ICLR 2024 · 105 citations
- X-Portrait: Expressive Portrait Animation with Hierarchical Motion AttentionYou Xie, Hongyi Xu, Guoxian Song, Chao Wang et al.SIGGRAPH 2024 · 40 citations
- MimicTalk: Mimicking a personalized and expressive 3D talking face in minutesZhenhui Ye, Tianyun Zhong, Yi Ren, Ziyue Jiang et al.NeurIPS 2024 · 28 citations
- Open-World Deepfake Attribution via Confidence-Aware Asymmetric LearningHaiyang Zheng, Nan Pu, Wenjing Li, Teng Long et al.AAAI 2026 · 5 citations
- FlashPortrait: 6x Faster Infinite Portrait Animation with Adaptive Latent PredictionShuyuan Tu, Yueming Pan, Yinming Huang, Xintong Han et al.CVPR 2026 · 3 citations
Builds on21
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Few-Shot Adversarial Learning of Realistic Neural Talking Head ModelsEgor Zakharov, Aliaksandra Shysheya, Egor Burkov, Victor S. LempitskyICCV 2019 · 687 citations
- Thin-Plate Spline Motion Model for Image AnimationJian Zhao, Hui ZhangCVPR 2022 · 196 citations
- MarioNETte: Few-Shot Face Reenactment Preserving Identity of Unseen TargetsSungjoo Ha, Martin Kersner, Beomsu Kim, Seokjun Seo et al.AAAI 2020 · 184 citations
Related papers
- Synergizing Motion and Appearance: Multi-Scale Compensatory Codebooks for Talking Head Video GenerationShuling Zhao, Fa-Ting Hong, Xiaoshui Huang, Dan XuCVPR 2025
- Occlusion-Insensitive Talking Head Video Generation via Facelet CompensationYuhui Deng, Yuqin Lu, Yangyang Xu, Yongwei Nie et al.AAAI 2025 · 3 citations
- FLNet: Landmark Driven Fetching and Learning Network for Faithful Talking Facial Animation SynthesisKuangxiao Gu, Yuqian Zhou, Thomas S. HuangAAAI 2020 · 63 citations
- EMMN: Emotional Motion Memory Network for Audio-driven Emotional Talking Face GenerationShuai Tan, Bin Ji, Ye PanICCV 2023 · 63 citations
- AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisYudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu et al.ICCV 2021 · 510 citations
