Cafe-Talk: Generating 3D Talking Face Animation with Multimodal Coarse- and Fine-grained Control
Hejia Chen, Haoxian Zhang, Shoulong Zhang, Xiaoqiang Liu, Sisi Zhuang, Yuan Zhang, Pengfei Wan, Di Zhang, Shuai Li
Abstract
3 Zhongguancun Laboratory ♣ Equal contribution ♠ Intern at Kuaishou Technology ♡ Corresponding author "Angry" & "A senior actor" "Laugh" … how to gh! <mute> Ha! Ha! Something like this? laucoarse ctrl speech render video fine ctrl language portrait label Figure 1: Adding multimodal coarse-and fine-grained control enables more flexible animations: Scenario: A senior actor is arguing with the director about how to smile. Action: The actor responds with anger and concludes with a sudden sarcastic laugh.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aca4b1bb-7b24-4345-a595-874f5aee1bcbCited by top-tier papers5
- OmniSync: Towards Universal Lip Synchronization via Diffusion TransformersZiqiao Peng, Jiwen Liu, Haoxian Zhang, Xiaoqiang Liu et al.NeurIPS 2025 · 30 citations
- EmoDiffTalk: Emotion-aware Diffusion for Editable 3D Gaussian Talking HeadChang Liu, Tianjiao Jing, Chengcheng Ma, Xuanqi Zhou et al.CVPR 2026 · 1 citation
- DualTalk: Dual-Speaker Interaction for 3D Talking Head ConversationsZiqiao Peng, Yanbo Fan, Haoyu Wu, Xuan Wang et al.CVPR 2025
- From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative BootstrappingXu He, Haoxian Zhang, Hejia Chen, Changyuan Zheng et al.ICML 2026
- A Focused Human Body Model for Accurate Anthropometric Measurements ExtractionShuhang Chen, Xianliang Huang, Zhizhou Zhong, Juhong Guan et al.CVPR 2025
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- MMHead: Towards Fine-grained Multi-modal 3D Facial AnimationSijing Wu, Yunhao Li, Yichao Yan, Huiyu Duan et al.ACM MM 2024 · 17 citations
- MoCha: Towards Movie-Grade Talking Character GenerationCong Wei, Bo Sun, Haoyu Ma, Ji Hou et al.NeurIPS 2025 · 2 citations
- MERMAID: Multi-perspective Self-reflective Agents with Generative Augmentation for Emotion RecognitionZhongyu Yang, Junhao Song, Siyang Song, Wei Pang et al.EMNLP 2025
- Sentiment and Emotion help Sarcasm? A Multi-task Learning Framework for Multi-Modal Sarcasm, Sentiment and Emotion AnalysisDushyant Singh Chauhan, Dhanush S. R, Asif Ekbal, Pushpak BhattacharyyaACL 2020 · 131 citations
- Predict and Use: Harnessing Predicted Gaze to Improve Multimodal Sarcasm DetectionDivyank Tiwari, Diptesh Kanojia, Anupama Ray, Apoorva Nunna et al.EMNLP 2023 · 12 citations
