MoCha: End-to-End Video Character Replacement without Structural Guidance
Zhengbo Xu, Jie Ma, Ziheng Wang, Zhan Peng, Jun Liang, Jing Li
Abstract
Controllable video character replacement with a user-provided identity remains a challenging problem due to the lack of paired video data. Prior works have predominantly relied on a reconstruction-based paradigm that requires per-frame segmentation masks and explicit structural guidance (e.g., skeleton, depth). This reliance, however, severely limits their generalizability in complex scenarios involving occlusions, character-object interactions, unusual poses, or challenging illumination, often leading to visual artifacts and temporal inconsistencies. In this paper, we propose MoCha, a pioneering framework that bypasses these limitations by requiring only a single arbitrary frame mask. To effectively adapt the multi-modal input condition and enhance facial identity, we introduce a condition-aware RoPE and employ an RL-based post-training stage. Furthermore, to overcome the scarcity of qualified paired-training data, we propose a comprehensive data construction pipeline. Specifically, we design three specialized datasets: a high-fidelity rendered dataset built with Unreal Engine 5 (UE5), an expression-driven dataset synthesized by current portrait animation techniques, and an augmented dataset derived from existing video-mask pairs. Extensive experiments demonstrate that our method substantially outperforms existing state-of-the-art approaches. We will release the code to facilitate further research. Please refer to our project page for more details: orange-3dv-team.github.io/MoCha
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a4c5dad7-6962-4c6d-b920-5ffed5dea0f2Builds on13
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- YOLOv12: Attention-Centric Real-Time Object DetectorsYunjie Tian, Qixiang Ye, David S. DoermannNeurIPS 2025 · 2,652 citations
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 1,105 citations
Related papers
- MoCha: Towards Movie-Grade Talking Character GenerationCong Wei, Bo Sun, Haoyu Ma, Ji Hou et al.NeurIPS 2025 · 2 citations
- MIMO: Controllable Character Video Synthesis with Spatial Decomposed ModelingYifang Men, Yuan Yao, Miaomiao Cui, Liefeng BoCVPR 2025
- TopoCap: Learning Topology-Agnostic Motion Priors for Monocular Video-to-AnimationCheng-Feng Pu, Jia-Peng Zhang, Meng-Hao Guo, Yan-Pei Cao et al.SIGGRAPH 2026
- ReactID: Synchronizing Realistic Actions and Identity in Personalized Video GenerationWei Li, Yiheng Zhang, Fuchen Long, Zhaofan Qiu et al.ICLR 2026
- Bringing Your Portrait to 3D PresenceJiawei Zhang, Lei Chu, Jiahao Li, Zhenyu Zang et al.CVPR 2026 · 3 citations
