Compositional 3D-aware Video Generation with LLM Director
Hanxin Zhu, Tianyu He, Anni Tang, Junliang Guo, Zhibo Chen, Jiang Bian
Abstract
Significant progress has been made in text-to-video generation through the use of powerful generative models and large-scale internet data. However, substantial challenges remain in precisely controlling individual concepts within the generated video, such as the motion and appearance of specific characters and the movement of viewpoints. In this work, we propose a novel paradigm that generates each concept in 3D representation separately and then composes them with priors from Large Language Models (LLM) and 2D diffusion models. Specifically, given an input textual prompt, our scheme consists of three stages: 1) We leverage LLM as the director to first decompose the complex query into several sub-prompts that indicate individual concepts within the video (e.g., scene, objects, motions), then we let LLM to invoke pre-trained expert models to obtain corresponding 3D representations of concepts. 2) To compose these representations, we prompt multi-modal LLM to produce coarse guidance on the scales and coordinates of trajectories for the objects. 3) To make the generated frames adhere to natural image distribution, we further leverage 2D diffusion priors and use Score Distillation Sampling to refine the composition. Extensive experiments demonstrate that our method can generate high-fidelity videos from text with diverse motion and flexible control over each concept. Project page: https://aka.ms/c3v.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a9242e0d-2ce2-430c-8fc0-bd0b931fdbe7Cited by top-tier papers7
- Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World ModelingHaoyu Wu, Diankun Wu, Tianyu He, Junliang Guo et al.ICLR 2026 · 89 citations
- Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-DistillationSherwin Bahmani, Tianchang Shen, Jiawei Ren, Jiahui Huang et al.ICLR 2026 · 33 citations
- Multi-Object Sketch Animation by Scene Decomposition and Motion PlanningJingyu Liu, Zijie Xin, Yuhan Fu, Ruixiang Zhao et al.ICCV 2025 · 6 citations
- AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion TransformersSherwin Bahmani, Ivan Skorokhodov, Guocheng Qian, Aliaksandr Siarohin et al.CVPR 2025
- Logical Guidance for the Exact Composition of Diffusion ModelsFrancesco Alesiani, Jonathan Warrell, Tanja Bien, Henrik Christiansen et al.ICML 2026
Builds on30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- RealisMotion: Decomposed Human Motion Control and Video Generation in the World SpaceJingyun Liang, Jingkai Zhou, Shikai Li, Chenjie Cao et al.ICML 2026 · 9 citations
- Modular-Cam: Modular Dynamic Camera-view Video Generation with LLMZirui Pan, Xin Wang, Yipeng Zhang, Hong Chen et al.AAAI 2025 · 6 citations
- LAMP: Language-Assisted Motion Planning for Controllable Video GenerationMuhammed Burak Kizil, Enes Şanlı, Niloy J. Mitra, Erkut Erdem et al.CVPR 2026 · 4 citations
- LLM-grounded Video Diffusion ModelsLong Lian, Baifeng Shi, Adam Yala, Trevor Darrell et al.ICLR 2024 · 87 citations
- BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video RepresentationsWeixi Feng, Chao Liu, Sifei Liu, William Yang Wang et al.CVPR 2025
