MultiShotMaster: A Controllable Multi-Shot Video Generation Framework
Qinghe Wang, Xiaoyu Shi, Baolu Li, Weikang Bian, Quande Liu, Huchuan Lu, Xintao Wang, Pengfei Wan, Kun Gai, Xu Jia
Abstract
Current video generation techniques excel at single-shot clips but struggle to produce narrative multi-shot videos, which require flexible shot arrangement, coherent narrative, and controllability beyond text prompts. To tackle these challenges, we propose MultiShotMaster, a framework for highly controllable multi-shot video generation. We extend a pretrained single-shot model by integrating two novel variants of RoPE. First, we introduce Multi-Shot Narrative RoPE, which applies explicit phase shift at shot transitions, enabling flexible shot arrangement while preserving the temporal narrative order. Second, we design Spatiotemporal Position-Aware RoPE to incorporate reference tokens and grounding signals, enabling spatiotemporal-grounded reference injection. In addition, to overcome data scarcity, we establish an automated data annotation pipeline to extract multi-shot videos, captions, cross-shot grounding signals and reference images. Our framework leverages the intrinsic architectural properties to support multi-shot video generation, featuring text-driven inter-shot consistency, customized subject with motion control, and background-driven customized scene. Both shot count and duration are flexibly configurable. Extensive experiments demonstrate the superior performance and outstanding controllability of our framework.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 461e7d88-0c40-4a91-9446-05a179844083Cited by top-tier papers2
- Group Editing: Edit Multiple Images in One GoYue Ma, Xinyu Wang, Qianli Ma, Qinghe Wang et al.CVPR 2026 · 15 citations
- Mode Seeking meets Mean Seeking for Fast Long Video GenerationShengqu Cai, Weili Nie, Chao Liu, Julius Berner et al.ICML 2026 · 9 citations
Builds on40
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- MV-S2V: Multi-View Subject-Consistent Video GenerationZiyang Song, Xinyu Gong, Bangya Liu, Zelin ZhaoSIGGRAPH 2026 · 1 citation
- ShotDirector: Directorially Controllable Multi-Shot Video Generation with Cinematographic TransitionsXiaoxue Wu, Xinyuan Chen, Yaohui Wang, Yu QiaoCVPR 2026 · 5 citations
- ReDirector: Creating Any-Length Video Retakes with Rotary Camera EncodingByeongjun Park, Byung-Hoon Kim, Hyungjin Chung, Jong ChulCVPR 2026 · 10 citations
- STAGE: Storyboard-Anchored Generation for Cinematic Multi-shot NarrativePeixuan Zhang, Zijian Jia, Kaiqi Liu, Shuchen Weng et al.CVPR 2026 · 25 citations
- OneStory: Coherent Multi-Shot Video Generation with Adaptive MemoryZhaochong An, Menglin Jia, Haonan Qiu, Zijian Zhou et al.CVPR 2026 · 33 citations
