Towards Realistic Visual Dubbing with Heterogeneous Sources
Tianyi Xie, Liucheng Liao, Cheng Bi, Benlai Tang, Xiang Yin, Jianfei Yang, Mingjie Wang, Jiali Yao, Yang Zhang, Zejun Ma
摘要
The task of few-shot visual dubbing focuses on synchronizing the lip movements with arbitrary speech input for any talking head video. Albeit moderate improvements in current approaches, they commonly require high-quality homologous data sources of videos and audios, thus causing the failure to leverage heterogeneous data sufficiently. In practice, it may be intractable to collect the perfect homologous data in some cases, for example, audio-corrupted or picture-blurry videos. To explore this kind of data and support high-fidelity few-shot visual dubbing, in this paper, we novelly propose a simple yet efficient two-stage framework with a higher flexibility of mining heterogeneous data. Specifically, our two-stage paradigm employs facial landmarks as intermediate prior of latent representations and disentangles the lip movements prediction from the core task of realistic talking head generation. By this means, our method makes it possible to independently utilize the training corpus for two-stage sub-networks using more available heterogeneous data easily acquired. Besides, thanks to the disentanglement, our framework allows a further fine-tuning for a given talking head, thereby leading to better speaker-identity preserving in the final synthesized results. Moreover, the proposed method can also transfer appearance features from others to the target speaker. Extensive experimental results demonstrate the superiority of our proposed method in generating highly realistic videos synchronized with the speech over the state-of-the-art.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- DINet: Deformation Inpainting Network for Realistic Face Visually Dubbing on High Resolution VideoZhimeng Zhang, Zhipeng Hu, Wenjin Deng, Changjie Fan 等AAAI 2023 · 被引用 106 次
- Adaptive Affine Transformation: A Simple and Effective Operation for Spatial Misaligned Image GenerationZhimeng Zhang, Yu DingACM MM 2022 · 被引用 18 次
- Hierarchically Controlled Deformable 3D Gaussians for Talking Head SynthesisZhenhua Wu, Linxuan Jiang, Xiang Li, Chaowei Fang 等AAAI 2025 · 被引用 2 次
- Learning to Dub Movies via Hierarchical Prosody ModelsGaoxiang Cong, Liang Li, Yuankai Qi, Zheng-Jun Zha 等CVPR 2023
- Monocular and Generalizable Gaussian Talking Head AnimationShengjie Gong, Haojie Li, Jiapeng Tang, Dongming Hu 等CVPR 2025
它引用的顶会 Paper4
- Free-Form Image Inpainting With Gated ConvolutionJiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen 等ICCV 2019 · 被引用 1,990 次
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 被引用 869 次
- Few-Shot Adversarial Learning of Realistic Neural Talking Head ModelsEgor Zakharov, Aliaksandra Shysheya, Egor Burkov, Victor S. LempitskyICCV 2019 · 被引用 687 次
- MarioNETte: Few-Shot Face Reenactment Preserving Identity of Unseen TargetsSungjoo Ha, Martin Kersner, Beomsu Kim, Seokjun Seo 等AAAI 2020 · 被引用 184 次
相关 Paper
- Learned Spatial Representations for Few-shot Talking-Head SynthesisMoustafa Meshry, Saksham Suri, Larry S. Davis, Abhinav ShrivastavaICCV 2021 · 被引用 51 次
- Talking Head Generation with Probabilistic Audio-to-Visual Diffusion PriorsZhentao Yu, Zixin Yin, Deyu Zhou, Duomin Wang 等ICCV 2023 · 被引用 65 次
- Identity-Preserving Talking Face Generation with Landmark and Appearance PriorsWeizhi Zhong, Chaowei Fang, Yinqi Cai, Pengxu Wei 等CVPR 2023
- From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative BootstrappingXu He, Haoxian Zhang, Hejia Chen, Changyuan Zheng 等ICML 2026
- AE-NeRF: Audio Enhanced Neural Radiance Field for Few Shot Talking Head SynthesisDongze Li, Kang Zhao, Wei Wang, Bo Peng 等AAAI 2024 · 被引用 25 次
