Learning Motion Refinement for Unsupervised Face Animation
Jiale Tao, Shuhang Gu, Wen Li, Lixin Duan
Abstract
Unsupervised face animation aims to generate a human face video based on the appearance of a source image, mimicking the motion from a driving video. Existing methods typically adopted a prior-based motion model (e.g., the local affine motion model or the local thin-plate-spline motion model). While it is able to capture the coarse facial motion, artifacts can often be observed around the tiny motion in local areas (e.g., lips and eyes), due to the limited ability of these methods to model the finer facial motions. In this work, we design a new unsupervised face animation approach to learn simultaneously the coarse and finer motions. In particular, while exploiting the local affine motion model to learn the global coarse facial motion, we design a novel motion refinement module to compensate for the local affine motion model for modeling finer face motions in local areas. The motion refinement is learned from the dense correlation between the source and driving images. Specifically, we first construct a structure correlation volume based on the keypoint features of the source and driving images. Then, we train a model to generate the tiny facial motions iteratively from low to high resolution. The learned motion refinements are combined with the coarse motion to generate the new image. Extensive experiments on widely used benchmarks demonstrate that our method achieves the best results among state-of-the-art baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bed73a10-ba18-46a8-97f8-c9c7496576d2Cited by top-tier papers4
- Occlusion-Insensitive Talking Head Video Generation via Facelet CompensationYuhui Deng, Yuqin Lu, Yangyang Xu, Yongwei Nie et al.AAAI 2025 · 3 citations
- Large Displacement Motion Transfer with Unsupervised Anytime InterpolationGuixiang Wang, Jianjun LiICML 2025
- Free-viewpoint Human Animation with Pose-correlated Reference SelectionFa-Ting Hong, Zhan Xu, Haiyang Liu, Qinjie Lin et al.CVPR 2025
- Synergizing Motion and Appearance: Multi-Scale Compensatory Codebooks for Talking Head Video GenerationShuling Zhao, Fa-Ting Hong, Xiaoshui Huang, Dan XuCVPR 2025
Builds on14
- FSGAN: Subject Agnostic Face Swapping and ReenactmentYuval Nirkin, Yosi Keller, Tal HassnerICCV 2019 · 710 citations
- Latent Image Animator: Learning to Animate Images via Latent Space NavigationYaohui Wang, Di Yang, François Brémond, Antitza DantchevaICLR 2022 · 219 citations
- Thin-Plate Spline Motion Model for Image AnimationJian Zhao, Hui ZhangCVPR 2022 · 196 citations
- MarioNETte: Few-Shot Face Reenactment Preserving Identity of Unseen TargetsSungjoo Ha, Martin Kersner, Beomsu Kim, Seokjun Seo et al.AAAI 2020 · 184 citations
- Depth-Aware Generative Adversarial Network for Talking Head Video GenerationFa-Ting Hong, Longhao Zhang, Li Shen, Dan XuCVPR 2022 · 168 citations
Related papers
- SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face AnimationWenxuan Zhang, Xiaodong Cun, Xuan Wang, Yong Zhang et al.CVPR 2023
- Motion Representations for Articulated AnimationAliaksandr Siarohin, Oliver J. Woodford, Jian Ren, Menglei Chai et al.CVPR 2021
- Structure-Aware Motion Transfer with Deformable Anchor ModelJiale Tao, Biao Wang, Borun Xu, Tiezheng Ge et al.CVPR 2022 · 33 citations
- Continuous Piecewise-Affine Based Motion Model for Image AnimationHexiang Wang, Fengqi Liu, Qianyu Zhou, Ran Yi et al.AAAI 2024 · 11 citations
- FG-Portrait: 3D Flow Guided Editable Portrait AnimationYating Xu, Yunqi Miao, Evangelos Ververas, Jiankang Deng et al.CVPR 2026
