Say Anything with Any Style
Shuai Tan, Bin Ji, Yu Ding, Ye Pan
摘要
Generating stylized talking head with diverse head motions is crucial for achieving natural-looking videos but still remains challenging. Previous works either adopt a regressive method to capture the speaking style, resulting in a coarse style that is averaged across all training data, or employ a universal network to synthesize videos with different styles which causes suboptimal performance. To address these, we propose a novel dynamic-weight method, namely Say Anything with Any Style (SAAS), which queries the discrete style representation via a generative model with a learned style codebook. Specifically, we develop a multi-task VQ-VAE that incorporates three closely related tasks to learn a style codebook as a prior for style extraction. This discrete prior, along with the generative model, enhances the precision and robustness when extracting the speaking styles of the given style clips. By utilizing the extracted style, a residual architecture comprising a canonical branch and style-specific branch is employed to predict the mouth shapes conditioned on any driving audio while transferring the speaking style from the source to any desired one. To adapt to different speaking styles, we steer clear of employing a universal network by exploring an elaborate HyperStyle to produce the style-specific weights offset for the style branch. Furthermore, we construct a pose generator and a pose codebook to store the quantized pose representation, allowing us to sample diverse head motions aligned with the audio and the extracted style. Experiments demonstrate that our approach surpasses state-of-the-art methods in terms of both lip-synchronization and stylized expression. Besides, we extend our SAAS to video-driven style editing field and achieve satisfactory performance as well.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- FlowVQTalker: High-Quality Emotional Talking Face Generation through Normalizing Flow and QuantizationShuai Tan, Bin Ji, Ye PanCVPR 2024 · 被引用 17 次
- DEEPTalk: Dynamic Emotion Embedding for Probabilistic Speech-Driven 3D Face AnimationJisoo Kim, Jungbin Cho, Joonho Park, Soonmin Hwang 等AAAI 2025 · 被引用 13 次
- SynMotion: Semantic-Visual Adaptation for Motion Customized Video GenerationShuai Tan, Biao Gong, Yujie Wei, Shiwei Zhang 等CVPR 2026 · 被引用 9 次
- EmoTaG: Emotion-Aware Talking Head Synthesis on Gaussian Splatting with Few-Shot PersonalizationHaolan Xu, Keli Cheng, Lei Wang, Ning Bi 等CVPR 2026 · 被引用 5 次
- MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled EmbeddingChang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja 等ACM MM 2025 · 被引用 3 次
它引用的顶会 Paper22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 被引用 869 次
- PIRenderer: Controllable Portrait Image Generation via Semantic Neural RenderingYurui Ren, Ge Li, Yuanqi Chen, Thomas H. Li 等ICCV 2021 · 被引用 284 次
- EAMM: One-Shot Emotional Talking Face via Audio-Based Emotion-Aware Motion ModelXinya Ji, Hang Zhou, Kaisiyuan Wang, Qianyi Wu 等SIGGRAPH 2022 · 被引用 150 次
- One-Shot Talking Face Generation from Single-Speaker Audio-Visual Correlation LearningSuzhen Wang, Lincheng Li, Yu Ding, Xin YuAAAI 2022 · 被引用 142 次
相关 Paper
- StyleTalk: One-Shot Talking Head Generation with Controllable Speaking StylesYifeng Ma, Suzhen Wang, Zhipeng Hu, Changjie Fan 等AAAI 2023 · 被引用 135 次
- SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face AnimationWenxuan Zhang, Xiaodong Cun, Xuan Wang, Yong Zhang 等CVPR 2023
- Synergizing Motion and Appearance: Multi-Scale Compensatory Codebooks for Talking Head Video GenerationShuling Zhao, Fa-Ting Hong, Xiaoshui Huang, Dan XuCVPR 2025
- VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeSicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang 等NeurIPS 2024 · 被引用 253 次
- GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific AdaptationWentao Hu, Shunkai Li, Ziqiao Peng, Haoxian Zhang 等ICCV 2025
