MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling
Yifang Men, Yuan Yao, Miaomiao Cui, Liefeng Bo
Abstract
Abstract Character video synthesis aims to produce realistic videos of animatable characters within lifelike scenes. As a fundamental problem in the computer vision and graphics community, 3D works typically require multi-view captures for per-case training, which severely limits their applicability of modeling arbitrary characters in a short time. Recent 2D methods break this limitation via pre-trained diffusion models, but they struggle for flexible controls, pose generality and scene interaction. To this end, we propose MIMO, a novel framework which can not only synthesize realistic character videos with controllable attributes (i.e., character, motion and scene) provided by simple user in-puts, but also simultaneously achieve advanced scalability to arbitrary characters, generality to novel 3D motions, and applicability to interactive real-world scenes in a unified framework. The core idea is to encode the 2D video to compact spatial codes, considering the inherent 3D nature of video occurrence. Concretely, we lift the 2D frame pixels into 3D using monocular depth estimators, and decompose the video clip into three spatial components (i.e., main human, underlying scene, and floating occlusion) in hierarchical layers based on the 3D depth. These components are further encoded to canonical identity code, structured motion code and full scene code, which are utilized as control signals of the synthesis process. The design of spatial decomposed modeling enables flexible user control, com-This CVPR paper is the Open Access version, provided by the Computer Vision Foundation. Except for this watermark, it is identical to the accepted version; the final published version of the proceedings is available on IEEE Xplore. plex motion expression, as well as 3D-aware synthesis for scene interactions. Experimental results show that the proposed method outperforms prior works by a large margin in character animation synthesis and is effective in providing a high degree of controllability (i.e., arbitrary characters, novel 3D motions, interactive scenes), thus enabling brandnew editing tasks (e.g., video character replacement).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d3154883-aa04-41a0-a9b4-b6493de047a8Cited by top-tier papers4
- PERSONA: Personalized Whole-Body 3D Avatar with Pose-Driven Deformations from a Single ImageGeonhee Sim, Gyeongsik MoonICCV 2025 · 6 citations
- Zero-Shot Reconstruction of Animatable 3D Avatars with Cloth Dynamics from a Single ImageJooHyun Kwon, Geonhee Sim, Gyeongsik MoonCVPR 2026 · 3 citations
- MotionWeaver: Holistic 4D-Anchored Framework for Multi-Humanoid Image AnimationXirui Hu, Yanbo Ding, Jiahao Wang, Tingting Shi et al.ICLR 2026 · 3 citations
- Video Camera Trajectory Editing with Generative Rendering from Estimated GeometryJunyoung Seo, Jisang Han, Jaewoo Jung, Siyoon Jin et al.AAAI 2026
Builds on37
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
Related papers
- Training-free Motion Factorization for Compositional Video GenerationZixuan Wang, Ziqin Zhou, Feng Chen, Duo Peng et al.CVPR 2026 · 1 citation
- RealisMotion: Decomposed Human Motion Control and Video Generation in the World SpaceJingyun Liang, Jingkai Zhou, Shikai Li, Chenjie Cao et al.ICML 2026 · 9 citations
- 3D-Aware Implicit Motion Control for View-Adaptive Human Video GenerationZhixue Fang, Xu He, Songlin Tang, Haoxian Zhang et al.CVPR 2026 · 4 citations
- Multi-Identity Human Image Animation with Structural Video DiffusionZhenzhi Wang, Yixuan Li, Yanhong Zeng, Yuwei Guo et al.ICCV 2025 · 1 citation
- Motion Synthesis with Sparse and Flexible Keyjoint ControlInwoo Hwang, Jinseok Bae, Donggeun Lim, Young Min KimICCV 2025 · 2 citations
