DINet: Deformation Inpainting Network for Realistic Face Visually Dubbing on High Resolution Video
Zhimeng Zhang, Zhipeng Hu, Wenjin Deng, Changjie Fan, Tangjie Lv, Yu Ding
摘要
For few-shot learning, it is still a critical challenge to realize photo-realistic face visually dubbing on high-resolution videos. Previous works fail to generate high-fidelity dubbing results. To address the above problem, this paper proposes a Deformation Inpainting Network (DINet) for high-resolution face visually dubbing. Different from previous works relying on multiple up-sample layers to directly generate pixels from latent embeddings, DINet performs spatial deformation on feature maps of reference images to better preserve highfrequency textural details. Specifically, DINet consists of one deformation part and one inpainting part. In the first part, five reference facial images adaptively perform spatial deformation to create deformed feature maps encoding mouth shapes at each frame, in order to align with the input driving audio and also the head poses of the input source images. In the second part, to produce face visually dubbing, a feature decoder is responsible for adaptively incorporating mouth movements from the deformed feature maps and other attributes (i.e., head pose and upper facial expression) from the source feature maps together. Finally, DINet achieves face visually dubbing with rich textural details. We conduct qualitative and quantitative comparisons to validate our DINet on high-resolution videos. The experimental results show that our method outperforms state-of-the-art works.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- SyncTalk: The Devil is in the Synchronization for Talking Head SynthesisZiqiao Peng, Wentao Hu, Yue Shi, Xiangyu Zhu 等CVPR 2024 · 被引用 65 次
- Say Anything with Any StyleShuai Tan, Bin Ji, Yu Ding, Ye PanAAAI 2024 · 被引用 30 次
- OmniSync: Towards Universal Lip Synchronization via Diffusion TransformersZiqiao Peng, Jiwen Liu, Haoxian Zhang, Xiaoqiang Liu 等NeurIPS 2025 · 被引用 30 次
- AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion EncodingTao Liu, Feilong Chen, Shuai Fan, Chenpeng Du 等ACM MM 2024 · 被引用 19 次
- PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head SynthesisYifan Xie, Tao Feng, Xin Zhang, Xiangyang Luo 等AAAI 2025 · 被引用 14 次
它引用的顶会 Paper16
- AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisYudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu 等ICCV 2021 · 被引用 510 次
- Latent Image Animator: Learning to Animate Images via Latent Space NavigationYaohui Wang, Di Yang, François Brémond, Antitza DantchevaICLR 2022 · 被引用 219 次
- HeadGAN: One-shot Neural Head Synthesis and EditingMichail Christos Doukas, Stefanos Zafeiriou, Viktoriia SharmanskaICCV 2021 · 被引用 164 次
- EAMM: One-Shot Emotional Talking Face via Audio-Based Emotion-Aware Motion ModelXinya Ji, Hang Zhou, Kaisiyuan Wang, Qianyi Wu 等SIGGRAPH 2022 · 被引用 150 次
- Expressive Talking Head Generation with Granular Audio-Visual ControlBorong Liang, Yan Pan, Zhizhi Guo, Hang Zhou 等CVPR 2022 · 被引用 114 次
相关 Paper
- Towards Realistic Visual Dubbing with Heterogeneous SourcesTianyi Xie, Liucheng Liao, Cheng Bi, Benlai Tang 等ACM MM 2021 · 被引用 34 次
- From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative BootstrappingXu He, Haoxian Zhang, Hejia Chen, Changyuan Zheng 等ICML 2026
- SIDGAN: High-Resolution Dubbed Video Generation via Shift-Invariant LearningUrwa Muaz, Wondong Jang, Rohun Tripathi, Santhosh Mani 等ICCV 2023 · 被引用 8 次
- SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face AnimationWenxuan Zhang, Xiaodong Cun, Xuan Wang, Yong Zhang 等CVPR 2023
- TACR-Net: Editing on Deep Video and Voice PortraitsLuchuan Song, Bin Liu, Guojun Yin, Xiaoyi Dong 等ACM MM 2021 · 被引用 20 次
