LAMP: Language-Assisted Motion Planning for Controllable Video Generation
Muhammed Burak Kizil, Enes Şanlı, Niloy J. Mitra, Erkut Erdem, Aykut Erdem, Duygu Ceylan
Abstract
Recent advances in video generation have achieved remarkable progress in visual fidelity and controllability, enabling conditioning not only on text but also on structural layout and motion signals. Among these, motion control (i.e., specifying both object dynamics and camera trajectories) is particularly critical for directing complex, cinematic scenes, yet existing interfaces remain limited. To address this gap, we introduce LAMP that leverages large language models (LLMs) as motion planners to translate natural language descriptions into explicit 3D trajectories for both dynamic objects and (relatively defined) cameras. Specifically, we fine-tune an LLM to generate frame-wise 3D bounding-box trajectories for objects and, conditioned on these, produce corresponding 3D camera paths, which are then converted into generator-compatible 2D control signals. We enable this by constructing a large-scale paired datasets through a combination of procedurally generated text–trajectory pairs and augmented real video datasets with 3D annotations. Experiments demonstrate improved controllability and alignment with user intent compared to state-of-the-art alternatives, establishing the first framework for joint object–camera trajectory generation directly from natural language.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ade9c23d-c32f-41c9-9b7d-db6c856e9569Cited by top-tier papers1
Ask how each one uses itBuilds on16
- Director3D: Real-world Camera Trajectory and 3D Scene Generation from TextXinyang Li, Zhangyu Lai, Linning Xu, Yansong Qu et al.NeurIPS 2024 · 60 citations
- VACE: All-in-One Video Creation and EditingZeyinzi Jiang, Zhen Han, Chaojie Mao, Jingfeng Zhang et al.ICCV 2025 · 58 citations
- Recammaster: Camera-Controlled Generative Rendering From a Single VideoJianhong Bai, Menghan Xia, Xiao Fu, Xintao Wang et al.ICCV 2025 · 33 citations
- TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion ModelsMark Yu, Wenbo Hu, Jinbo Xing, Ying ShanICCV 2025 · 25 citations
- Stable Virtual Camera: Generative View Synthesis with Diffusion ModelsJensen Zhou, Hang Gao, Vikram Voleti, Aaryaman Vasishta et al.ICCV 2025 · 25 citations
Related papers
- ChatCam: Empowering Camera Control through Conversational AIXinhang Liu, Yu-Wing Tai, Chi-Keung TangNeurIPS 2024 · 19 citations
- LLM-grounded Video Diffusion ModelsLong Lian, Baifeng Shi, Adam Yala, Trevor Darrell et al.ICLR 2024 · 87 citations
- Compositional 3D-aware Video Generation with LLM DirectorHanxin Zhu, Tianyu He, Anni Tang, Junliang Guo et al.NeurIPS 2024 · 19 citations
- LaMP: Language-Motion Pretraining for Motion Generation, Retrieval, and CaptioningZhe Li, Weihao Yuan, Yisheng He, Lingteng Qiu et al.ICLR 2025
- Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional TokenizationYang Jin, Zhicheng Sun, Kun Xu, Kun Xu et al.ICML 2024 · 94 citations
