DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
Zun Wang, Jialu Li, Han Lin, Jaehong Yoon, Mohit Bansal
摘要
Storytelling video generation (SVG) aims to produce coherent and visually rich multi-scene videos that follow a structured narrative. Existing methods primarily employ LLM for high-level planning to decompose a story into scene-level descriptions, which are then independently generated and stitched together. However, these approaches struggle with generating high-quality videos aligned with the complex single-scene description, as visualizing such complex description involves coherent composition of multiple objects/events, complex motion synthesis and character customization with sequential motions. To address these challenges, we propose DREAMRUNNER, a novel story-to-video generation method: First, we structure the input script using a large language model (LLM) to facilitate both coarse-grained scene planning as well as fine-grained object-level layout planning. Next, DREAMRUNNER presents retrieval-augmented test-time adaptation to capture target motion priors for objects in each scene, supporting diverse motion customization based on retrieved videos, thus facilitating the generation of new videos with complex, scripted motions. Lastly, we propose a novel spatial-temporal region-based 3D attention and prior injection module SR3AI for fine-grained object-motion binding and frame-by-frame spatial-temporal semantic control. We compare DREAMRUNNER with various SVG baselines, demonstrating state-of-the-art performance in character consistency, text alignment, and smooth transitions. Additionally, DREAMRUNNER exhibits strong fine-grained condition-following ability in compositional text-to-video generation, significantly outperforming baselines on T2V-ComBench. Finally, we demonstrate DREAMRUNNER’s ability to generate multi-character interactions with qualitative examples.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Long Context Tuning for Video GenerationYuwei Guo, Ceyuan Yang, Ziyan Yang, Zhibei Ma 等ICCV 2025 · 被引用 6 次
- TS-Attn: Temporal-wise Separable Attention for Multi-Event Video GenerationHongyu Zhang, Yufan Deng, Zilin Pan, Peng-Tao Jiang 等ICLR 2026 · 被引用 5 次
它引用的顶会 Paper38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- MUSIQ: Multi-scale Image Quality TransformerJunjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar 等ICCV 2021 · 被引用 1,325 次
- Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video GenerationJay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei 等ICCV 2023 · 被引用 1,113 次
相关 Paper
- MoTrans: Customized Motion Transfer with Text-driven Video Diffusion ModelsXiaomin Li, Xu Jia, Qinghe Wang, Haiwen Diao 等ACM MM 2024 · 被引用 6 次
- Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal ConditioningZhengjian Yao, Yongzhi Li, Xinyuan Gao, Quan Chen 等CVPR 2026 · 被引用 3 次
- Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video GenerationGuy Yariv, Yuval Kirstain, Amit Zohar, Shelly Sheynin 等CVPR 2025
- Comp-Attn: Present-and-Align Attention for Compositional Video GenerationHongyu Zhang, Yufan Deng, Shenghai Yuan, Xuehan Hou 等ICML 2026
- Human Motion Generation in 3D Scenes from Open-Ended Textual Instructions with MLLM PlanningSiyi Qian, Jian Fang, Yuzhou Mao, Yayun Zou 等ACM MM 2025
