PrevizWhiz: Combining Rough 3D Scenes and 2D Video to Guide Generative Video Previsualization
Erzhen Hu, Frederik Brudy, David Ledo, George W. Fitzmaurice, Fraser Anderson
Abstract
In pre-production, filmmakers and 3D animation experts must rapidly prototype ideas to explore a film’s possibilities before full-scale production, yet conventional approaches involve trade-offs in efficiency and expressiveness. Hand-drawn storyboards often lack spatial precision needed for complex cinematography, while 3D previsualization demands expertise and high-quality rigged assets. To address this gap, we present PrevizWhiz, a system that leverages rough 3D scenes in combination with generative image and video models to create stylized video previews. The workflow integrates frame-level image restyling with adjustable resemblance, time-based editing through motion paths or external video inputs, and refinement into high-fidelity video clips. A study with filmmakers demonstrates that our system lowers technical barriers for film-makers, accelerates creative iteration, and effectively bridges the communication gap, while also surfacing challenges of continuity, authorship, and ethical consideration in AI-assisted filmmaking.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd3c906f-99df-492c-a012-ea2affc94b65Builds on22
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- CollageVis: Rapid Previsualization Tool for Indie Filmmaking using Video CollagesHye-Young Jo, Ryo Suzuki, Yoonji KimCHI 2024 · 5 citations
- CineVision: An Interactive Pre-visualization Storyboard System for Director-Cinematographer CollaborationZheng Wei, Hongtao Wu, Lvmin Zhang, Xian Xu et al.UIST 2025 · 8 citations
- One Week in the Future: Previs Design Futuring for HCI ResearchAlexander Ivanov, Tim Au Yeung, Kathryn Blair, Kurtis Thorvald Danyluk et al.CHI 2022 · 11 citations
- SketchDynamics: Exploring Free-Form Sketches for Dynamic Intent Expression in Animation GenerationBoyu Li, Lin-Ping Yuan, Zeyu Wang, Hongbo FuCHI 2026 · 1 citation
- I2V3D: Controllable Image-to-Video Generation with 3D GuidanceZhiyuan Zhang, Dongdong Chen, Jing LiaoICCV 2025 · 3 citations
