Learned Spatial Representations for Few-shot Talking-Head Synthesis
Moustafa Meshry, Saksham Suri, Larry S. Davis, Abhinav Shrivastava
摘要
We propose a novel approach for few-shot talking-head synthesis. While recent works in neural talking heads have produced promising results, they can still produce images that do not preserve the identity of the subject in source images. We posit this is a result of the entangled representation of each subject in a single latent code that models 3D shape information, identity cues, colors, lighting and even background details. In contrast, we propose to factorize the representation of a subject into its spatial and style components. Our method generates a target frame in two steps. First, it predicts a discrete and dense spatial layout for the target image. Second, an image generator utilizes the predicted layout for spatial denormalization and synthesizes the target frame. We experimentally show that this disentangled representation leads to a significant improvement over previous methods, both quantitatively and qualitatively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- SyncTalk: The Devil is in the Synchronization for Talking Head SynthesisZiqiao Peng, Wentao Hu, Yue Shi, Xiangyu Zhu 等CVPR 2024 · 被引用 65 次
- HyperReenact: One-Shot Reenactment via Jointly Learning to Refine and Retarget FacesStella Bounareli, Christos Tzelepis, Vasileios Argyriou, Ioannis Patras 等ICCV 2023 · 被引用 63 次
- Talking Head from Speech Audio using a Pre-trained Image GeneratorMohammed M. Alghamdi, He Wang, Andrew J. Bulpitt, David C. HoggACM MM 2022 · 被引用 25 次
- SegTalker: Segmentation-based Talking Face Generation with Mask-guided Local EditingLingyu Xiong, Xize Cheng, Jintao Tan, Xianjia Wu 等ACM MM 2024 · 被引用 10 次
- SIDGAN: High-Resolution Dubbed Video Generation via Shift-Invariant LearningUrwa Muaz, Wondong Jang, Rohun Tripathi, Santhosh Mani 等ICCV 2023 · 被引用 8 次
它引用的顶会 Paper10
- FSGAN: Subject Agnostic Face Swapping and ReenactmentYuval Nirkin, Yosi Keller, Tal HassnerICCV 2019 · 被引用 710 次
- Few-Shot Adversarial Learning of Realistic Neural Talking Head ModelsEgor Zakharov, Aliaksandra Shysheya, Egor Burkov, Victor S. LempitskyICCV 2019 · 被引用 687 次
- MarioNETte: Few-Shot Face Reenactment Preserving Identity of Unseen TargetsSungjoo Ha, Martin Kersner, Beomsu Kim, Seokjun Seo 等AAAI 2020 · 被引用 184 次
- FLNet: Landmark Driven Fetching and Learning Network for Faithful Talking Facial Animation SynthesisKuangxiao Gu, Yuqian Zhou, Thomas S. HuangAAAI 2020 · 被引用 63 次
- One-Shot Free-View Neural Talking-Head Synthesis for Video ConferencingTing-Chun Wang, Arun Mallya, Ming-Yu LiuCVPR 2021
相关 Paper
- Towards Realistic Visual Dubbing with Heterogeneous SourcesTianyi Xie, Liucheng Liao, Cheng Bi, Benlai Tang 等ACM MM 2021 · 被引用 34 次
- That's What I Said: Fully-Controllable Talking Face GenerationYoungjoon Jang, Kyeongha Rho, Jong-Bin Woo, Hyeongkeun Lee 等ACM MM 2023 · 被引用 7 次
- StyleTalk: One-Shot Talking Head Generation with Controllable Speaking StylesYifeng Ma, Suzhen Wang, Zhipeng Hu, Changjie Fan 等AAAI 2023 · 被引用 135 次
- AE-NeRF: Audio Enhanced Neural Radiance Field for Few Shot Talking Head SynthesisDongze Li, Kang Zhao, Wei Wang, Bo Peng 等AAAI 2024 · 被引用 25 次
- DiffTalk: Crafting Diffusion Models for Generalized Audio-Driven Portraits AnimationShuai Shen, Wenliang Zhao, Zibin Meng, Wanhua Li 等CVPR 2023
