MoCa: Modeling Object Consistency for 3D Camera Control in Video Generation
Zhijing Cheng, Xuancheng Zhang, Donglin Di, Chen Wei, Hao Li, Xun Yang
Abstract
Camera control is important in text-to-video generation for achieving realistic scene navigation and view synthesis. This control is defined by parameters that describe movement through 3D space, thereby introducing 3D consistency into the generation process. A core challenge for existing methods is achieving 3D consistency within the 2D pixel domain. Strategies that directly integrate camera conditions into text-to-video models often produce artifacts, while those relying on explicit 3D supervision face challenges with generalization. Both limitations originate from the gap between the 2D pixel space and the underlying 3D world. The key insight is that the projection of a smooth 3D camera movement produces consistency in object view, appearance, and motion across 2D frames. Inspired by this insight, we propose MoCa, a dual-branch framework that bridges this gap by modeling object consistency to implicitly learn 3D relationships between the camera and the scene. To ensure view consistency, we design a Spatial-Temporal Camera Encoder with Plücker embedding, which encodes camera trajectories into a geometrically grounded latent representation. For appearance consistency, we introduce a semantic guidance strategy that leverages persistent vision-language features to maintain object identity and texture across frames. To address motion consistency, we propose an object-aware motion disentanglement mechanism that separates object dynamics from global camera movement, ensuring precise camera control and natural object motion. Experiments show that MoCa achieves accurate camera control while preserving video quality, offering a practical and effective solution for camera-controllable video generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c214a8bc-de23-4263-8b85-cb947a247e95Builds on22
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
- Light Field Networks: Neural Scene Representations with Single-Evaluation RenderingVincent Sitzmann, Semon Rezchikov, Bill Freeman, Josh Tenenbaum et al.NeurIPS 2021 · 426 citations
- Make-A-Video: Text-to-Video Generation without Text-Video DataUriel Singer, Adam Polyak, Thomas Hayes, Xi Yin et al.ICLR 2023 · 313 citations
- StyleGAN-V: A Continuous Video Generator with the Price, Image Quality and Perks of StyleGAN2Ivan Skorokhodov, Sergey Tulyakov, Mohamed ElhoseinyCVPR 2022 · 167 citations
Related papers
- VD3D: Taming Large Video Diffusion Transformers for 3D Camera ControlSherwin Bahmani, Ivan Skorokhodov, Aliaksandr Siarohin, Willi Menapace et al.ICLR 2025
- 3D-Aware Implicit Motion Control for View-Adaptive Human Video GenerationZhixue Fang, Xu He, Songlin Tang, Haoxian Zhang et al.CVPR 2026 · 4 citations
- Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated AttentionDejia Xu, Yifan Jiang, Chen Huang, Liangchen Song et al.ICML 2025
- Modular-Cam: Modular Dynamic Camera-view Video Generation with LLMZirui Pan, Xin Wang, Yipeng Zhang, Hong Chen et al.AAAI 2025 · 6 citations
- SceneScape: Text-Driven Consistent Scene GenerationRafail Fridman, Amit Abecasis, Yoni Kasten, Tali DekelNeurIPS 2023 · 196 citations
