FSRT: Facial Scene Representation Transformer for Face Reenactment from Factorized Appearance, Head-Pose, and Facial Expression Features
Andre Rochow, Max Schwarz, Sven Behnke
摘要
The task of face reenactment is to transfer the head motion and facial expressions from a driving video to the appearance of a source image, which may be of a different person (cross-reenactment). Most existing methods are CNN-based and estimate optical flow from the source image to the current driving frame, which is then inpainted and refined to produce the output animation. We propose a transformer-based encoder for computing a set-latent representation of the source image(s). We then predict the output color of a query pixel using a transformer-based decoder, which is conditioned with keypoints and a facial expression vector extracted from the driving frame. Latent representations of the source person are learned in a self-supervised manner that factorize their appearance, head pose, and facial expressions. Thus, they are perfectly suited for cross-reenactment. In contrast to most related work, our method naturally extends to multiple source images and can thus adapt to person-specific facial dynamics. We also propose data augmentation and regularization schemes that are necessary to prevent overfitting and support generalizability of the learned representations. We evaluated our approach in a randomized user study. The results indicate superior performance compared to the state-of-the-art in terms of motion transfer quality and temporal consistency.<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup><sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>Code & Video: https://andrerochow.github.io/fsrt
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Thinking Racial Bias in Fair Forgery Detection: Models, Datasets and EvaluationsDecheng Liu, Zongqi Wang, Chunlei Peng, Nannan Wang 等AAAI 2025 · 被引用 11 次
- PortraitDirector: A Hierarchical Disentanglement Framework for Controllable and Real-time Facial ReenactmentChaonan Ji, Jinwei Qi, Sheng Xu, Peng Zhang 等CVPR 2026
它引用的顶会 Paper29
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised LearningAdrien Bardes, Jean Ponce, Yann LeCunICLR 2022 · 被引用 1,226 次
- Few-Shot Adversarial Learning of Realistic Neural Talking Head ModelsEgor Zakharov, Aliaksandra Shysheya, Egor Burkov, Victor S. LempitskyICCV 2019 · 被引用 687 次
- SC-FEGAN: Face Editing Generative Adversarial Network With User's Sketch and ColorYoungjoo Jo, Jongyoul ParkICCV 2019 · 被引用 325 次
相关 Paper
- ToonTalker: Cross-Domain Face ReenactmentYuan Gong, Yong Zhang, Xiaodong Cun, Fei Yin 等ICCV 2023 · 被引用 14 次
- Neural Head Reenactment with Latent Pose DescriptorsEgor Burkov, Igor Pasechnik, Artur Grigorev, Victor S. LempitskyCVPR 2020
- Learning Identity-Invariant Motion Representations for Cross-ID Face ReenactmentPo-Hsiang Huang, Fu-En Yang, Yu-Chiang Frank WangCVPR 2020
- Mesh Guided One-shot Face Reenactment Using Graph Convolutional NetworksGuangming Yao, Yi Yuan, Tianjia Shao, Kun ZhouACM MM 2020 · 被引用 42 次
- LatentAvatar: Learning Latent Expression Code for Expressive Neural Head AvatarYuelang Xu, Hongwen Zhang, Lizhen Wang, Xiaochen Zhao 等SIGGRAPH 2023 · 被引用 40 次
