MoCha: Towards Movie-Grade Talking Character Generation
Cong Wei, Bo Sun, Haoyu Ma, Ji Hou, Felix Juefei-Xu, Zecheng He, Xiaoliang Dai, Luxin Zhang, Kunpeng Li, Tingbo Hou, Animesh Sinha, Peter Vajda, Wenhu Chen
摘要
"Two distinct streams of tears trail down her cheeks as she speaks with an angry expression…" "A medium shot of a man interacting warmly with an elephant. the man talks to the camera…" "A tilt up shot of a man standing in a dimly lit room, speaking to the camera…" Action Control Multi-Character Turn-based Talk Talking Character "Close-up shot of a doctor in a white lab coat over blue scrubs, speaking…" Emotion Control Figure 1: MoCha is an end-to-end dialogue-centric video generation model that takes only speech and text as input, without requiring any auxiliary conditions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper28
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang 等ICLR 2024 · 被引用 1,493 次
- Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video GenerationJay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei 等ICCV 2023 · 被引用 1,113 次
相关 Paper
- MoCha: End-to-End Video Character Replacement without Structural GuidanceZhengbo Xu, Jie Ma, Ziheng Wang, Zhan Peng 等CVPR 2026 · 被引用 9 次
- Mind the Time: Temporally-Controlled Multi-Event Video GenerationZiyi Wu, Aliaksandr Siarohin, Willi Menapace, Ivan Skorokhodov 等CVPR 2025
- Audio-Visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head GenerationFa-Ting Hong, Zunnan Xu, Zixiang Zhou, Jun Zhou 等ICCV 2025 · 被引用 2 次
- Cafe-Talk: Generating 3D Talking Face Animation with Multimodal Coarse- and Fine-grained ControlHejia Chen, Haoxian Zhang, Shoulong Zhang, Xiaoqiang Liu 等ICLR 2025
- MoCa: Modeling Object Consistency for 3D Camera Control in Video GenerationZhijing Cheng, Xuancheng Zhang, Donglin Di, Chen Wei 等ICLR 2026
