MoCha: Towards Movie-Grade Talking Character Generation
Cong Wei, Bo Sun, Haoyu Ma, Ji Hou, Felix Juefei-Xu, Zecheng He, Xiaoliang Dai, Luxin Zhang, Kunpeng Li, Tingbo Hou, Animesh Sinha, Peter Vajda, Wenhu Chen
Abstract
"Two distinct streams of tears trail down her cheeks as she speaks with an angry expression…" "A medium shot of a man interacting warmly with an elephant. the man talks to the camera…" "A tilt up shot of a man standing in a dimly lit room, speaking to the camera…" Action Control Multi-Character Turn-based Talk Talking Character "Close-up shot of a doctor in a white lab coat over blue scrubs, speaking…" Emotion Control Figure 1: MoCha is an end-to-end dialogue-centric video generation model that takes only speech and text as input, without requiring any auxiliary conditions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 880f56aa-937f-49dd-840d-b26c8c588ae8Cited by top-tier papers1
Ask how each one uses itBuilds on28
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
- Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video GenerationJay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei et al.ICCV 2023 · 1,113 citations
Related papers
- MoCha: End-to-End Video Character Replacement without Structural GuidanceZhengbo Xu, Jie Ma, Ziheng Wang, Zhan Peng et al.CVPR 2026 · 9 citations
- Mind the Time: Temporally-Controlled Multi-Event Video GenerationZiyi Wu, Aliaksandr Siarohin, Willi Menapace, Ivan Skorokhodov et al.CVPR 2025
- Audio-Visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head GenerationFa-Ting Hong, Zunnan Xu, Zixiang Zhou, Jun Zhou et al.ICCV 2025 · 2 citations
- Cafe-Talk: Generating 3D Talking Face Animation with Multimodal Coarse- and Fine-grained ControlHejia Chen, Haoxian Zhang, Shoulong Zhang, Xiaoqiang Liu et al.ICLR 2025
- MoCa: Modeling Object Consistency for 3D Camera Control in Video GenerationZhijing Cheng, Xuancheng Zhang, Donglin Di, Chen Wei et al.ICLR 2026
