Image Conductor: Precision Control for Interactive Video Synthesis
Yaowei Li, Xintao Wang, Zhaoyang Zhang, Zhouxia Wang, Ziyang Yuan, Liangbin Xie, Ying Shan, Yuexian Zou
Abstract
Filmmaking and animation production often require sophisticated techniques for coordinating camera transitions and object movements, typically involving labor-intensive real-world capturing. Despite advancements in generative AI for video creation, achieving precise control over motion for interactive video asset generation remains challenging. To this end, we propose Image Conductor, a method for precise control of camera transitions and object movements to generate video assets from a single image. An well-cultivated training strategy is proposed to separate distinct camera and object motion by camera LoRA weights and object LoRA weights. To further eliminate motion ambiguity from ill-posed trajectories, we introduce a camera-free guidance technique during inference process, enhancing object movements while eliminating camera transitions. Additionally, we develop a trajectory-oriented video motion data curation pipeline for training. Quantitative and qualitative experiments demonstrate our method's precision and fine-grained control in generating motion-controllable videos from images, advancing the practical application of interactive video synthesis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f37cfe67-8354-4f15-9756-0e7e7e1728f9Cited by top-tier papers34
- MotionStream: Real-Time Video Generation with Interactive Motion ControlsJoonghyuk Shin, Zhengqi Li, Richard Zhang, Jun-Yan Zhu et al.ICLR 2026 · 79 citations
- Wan-Move: Motion-controllable Video Generation via Latent Trajectory GuidanceRuihang Chu, Yefei He, Zhekai Chen, Shiwei Zhang et al.NeurIPS 2025 · 50 citations
- Physics-Driven Spatiotemporal Modeling for AI-Generated Video DetectionShuhai Zhang, Zihao Lian, Jiahao Yang, Daiyuan Li et al.NeurIPS 2025 · 29 citations
- VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric ControlSixiao Zheng, Minghao Yin, Wenbo Hu, Xiaoyu Li et al.CVPR 2026 · 27 citations
- Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation ControlZekai Gu, Rui Yan, Jiahao Lu, Peng Li et al.SIGGRAPH 2025 · 21 citations
Builds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
Related papers
- I2V3D: Controllable Image-to-Video Generation with 3D GuidanceZhiyuan Zhang, Dongdong Chen, Jing LiaoICCV 2025 · 3 citations
- Motion Modes: What Could Happen Next?Karran Pandey, Yannick Hold-Geoffroy, Matheus Gadelha, Niloy J. Mitra et al.CVPR 2025
- MotionCtrl: A Unified and Flexible Motion Controller for Video GenerationZhouxia Wang, Ziyang Yuan, Xintao Wang, Yaowei Li et al.SIGGRAPH 2024 · 123 citations
- COMD: Training-free Video Motion Transfer With Camera-Object Motion DisentanglementTeng Hu, Jiangning Zhang, Ran Yi, Yating Wang et al.ACM MM 2024 · 1 citation
- MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory GuidanceQuanhao Li, Zhen Xing, Rui Wang, Hui Zhang et al.ICCV 2025 · 10 citations
