Flow-Guided One-Shot Talking Face Generation With a High-Resolution Audio-Visual Dataset
Zhimeng Zhang, Lincheng Li, Yu Ding, Changjie Fan
Abstract
The new dataset is collected from youtube and consists of about 16 hours 720P or 1080P videos. We leverage the facial 3D morphable model (3DMM) to split the framework into two cascaded modules instead of learning a direct mapping from audio to video. In the first module, we propose a novel animation generator to produce the movements of mouth, eyebrow and head pose simultaneously. In the second module, we transform animation into dense flow to provide more expression details and carefully design a novel flow-guided video generator to synthesize videos. Our method is able to produce high-definition videos and outperforms state-of-the-art works in objective and subjective comparisons * .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4697f44b-9238-43ad-b9e2-a6cdf84d579eCited by top-tier papers160
- AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisYudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu et al.ICCV 2021 · 510 citations
- VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeSicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang et al.NeurIPS 2024 · 253 citations
- EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark ConditionsZhiyuan Chen, Jiajiong Cao, Zhiquan Chen, Yuming Li et al.AAAI 2025 · 197 citations
- EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face AnimationZiqiao Peng, Haoyu Wu, Zhenbo Song, Hao Xu et al.ICCV 2023 · 192 citations
- StyleTalk: One-Shot Talking Head Generation with Controllable Speaking StylesYifeng Ma, Suzhen Wang, Zhipeng Hu, Changjie Fan et al.AAAI 2023 · 135 citations
Builds on3
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- Write-a-speaker: Text-based Emotional and Rhythmic Talking-head GenerationLincheng Li, Suzhen Wang, Zhimeng Zhang, Yu Ding et al.AAAI 2021 · 88 citations
- FLNet: Landmark Driven Fetching and Learning Network for Faithful Talking Facial Animation SynthesisKuangxiao Gu, Yuqian Zhou, Thomas S. HuangAAAI 2020 · 63 citations
Related papers
- SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face AnimationWenxuan Zhang, Xiaodong Cun, Xuan Wang, Yong Zhang et al.CVPR 2023
- ECAvatar: 3D Avatar Facial Animation with Controllable Identity and EmotionMinjing Yu, Delong Pang, Ziwen Kang, Zhiyao Sun et al.ACM MM 2024 · 4 citations
- Imitating Arbitrary Talking Style for Realistic Audio-Driven Talking Face SynthesisHaozhe Wu, Jia Jia, Haoyu Wang, Yishun Dou et al.ACM MM 2021 · 50 citations
- EMMN: Emotional Motion Memory Network for Audio-driven Emotional Talking Face GenerationShuai Tan, Bin Ji, Ye PanICCV 2023 · 63 citations
- RigNeRF: Fully Controllable Neural 3D PortraitsShahRukh Athar, Zexiang Xu, Kalyan Sunkavalli, Eli Shechtman et al.CVPR 2022 · 117 citations
