Imitating Arbitrary Talking Style for Realistic Audio-Driven Talking Face Synthesis
Haozhe Wu, Jia Jia, Haoyu Wang, Yishun Dou, Chao Duan, Qingshan Deng
Abstract
People talk with diversified styles. For one piece of speech, different talking styles exhibit significant differences in the facial and head pose movements. For example, the "excited" style usually talks with the mouth wide open, while the "solemn" style is more standardized and seldomly exhibits exaggerated motions. Due to such huge differences between different styles, it is necessary to incorporate the talking style into audio-driven talking face synthesis framework. In this paper, we propose to inject style into the talking face synthesis framework through imitating arbitrary talking style of the particular reference video. Specifically, we systematically investigate talking styles with our collected Ted-HD dataset and construct style codes as several statistics of 3D morphable model (3DMM) parameters. Afterwards, we devise a latent-style-fusion (LSF) model to synthesize stylized talking faces by imitating talking styles from the style codes. We emphasize the following novel characteristics of our framework: (1) It doesn't require any annotation of the style, the talking style is learned in an unsupervised manner from talking videos in the wild. (2) It can imitate arbitrary styles from arbitrary videos, and the style codes can also be interpolated to generate new styles. Extensive experiments demonstrate that the proposed framework has the ability to synthesize more natural and expressive talking styles compared with baseline methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da49bb72-61d4-4160-a7d8-4e30dd41eb57Cited by top-tier papers19
- StyleTalk: One-Shot Talking Head Generation with Controllable Speaking StylesYifeng Ma, Suzhen Wang, Zhipeng Hu, Changjie Fan et al.AAAI 2023 · 135 citations
- DINet: Deformation Inpainting Network for Realistic Face Visually Dubbing on High Resolution VideoZhimeng Zhang, Zhipeng Hu, Wenjin Deng, Changjie Fan et al.AAAI 2023 · 106 citations
- Real3D-Portrait: One-shot Realistic 3D Talking Portrait SynthesisZhenhui Ye, Tianyun Zhong, Yi Ren, Jiaqi Yang et al.ICLR 2024 · 105 citations
- READ: Large-Scale Neural Scene Rendering for Autonomous DrivingZhuopeng Li, Lu Li, Jianke ZhuAAAI 2023 · 78 citations
- Implicit Identity Representation Conditioned Memory Compensation Network for Talking Head Video GenerationFa-Ting Hong, Dan XuICCV 2023 · 75 citations
Builds on3
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- Talking Face Generation with Expression-Tailored Generative Adversarial NetworkDan Zeng, Han Liu, Hui Lin, Shiming GeACM MM 2020 · 30 citations
- Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual RepresentationHang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy et al.CVPR 2021
Related papers
- DiffPoseTalk: Speech-Driven Stylistic 3D Facial Animation and Head Pose Generation via Diffusion ModelsZhiyao Sun, Tian Lv, Sheng Ye, Matthieu Gaetan Lin et al.SIGGRAPH 2024 · 81 citations
- Flow-Guided One-Shot Talking Face Generation With a High-Resolution Audio-Visual DatasetZhimeng Zhang, Lincheng Li, Yu Ding, Changjie FanCVPR 2021
- Style2Talker: High-Resolution Talking Head Generation with Emotion Style and Art StyleShuai Tan, Bin Ji, Ye PanAAAI 2024
- PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face GenerationBaiqin Wang, Xiangyu Zhu, Fan Shen, Hao Xu et al.CVPR 2026 · 8 citations
- High-Fidelity Generalized Emotional Talking Face Generation with Multi-Modal Emotion Space LearningChao Xu, Junwei Zhu, Jiangning Zhang, Yue Han et al.CVPR 2023
