Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-Distillation
Sherwin Bahmani, Tianchang Shen, Jiawei Ren, Jiahui Huang, Yifeng Jiang, Haithem Turki, Andrea Tagliasacchi, David B. Lindell, Zan Gojcic, Sanja Fidler, Huan Ling, Jun Gao, Xuanchi Ren
Abstract
The ability to generate virtual environments is crucial for applications ranging from gaming to physical AI domains such as robotics, autonomous driving, and industrial AI. Current learning-based 3D reconstruction methods rely on the availability of captured real-world multi-view data, which is not always readily available. Recent advancements in video diffusion models have shown remarkable imagination capabilities, yet their 2D nature prevents their use in simulations where a robot needs to navigate and interact with the environment. In this paper, we propose a self-distillation framework that aims to distill the implicit 3D knowledge in the video diffusion models into an explicit 3D Gaussian Splatting (3DGS) representation, eliminating the need for multi-view training data. Specifically, we augment the typical RGB decoder with a 3DGS decoder, which is supervised by the output of the RGB decoder. In this approach, the 3DGS decoder can be purely trained with synthetic data generated by video diffusion models. At inference time, our model can synthesize 3D scenes from either a text prompt or a single image for real-time rendering. Our framework further extends to dynamic 3D scene generation from a monocular input video. Experimental results show that our framework achieves state-of-the-art performance in static and dynamic 3D scene generation. Video results: https://research.nvidia.com/labs/toronto-ai/lyra
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c9c9f986-2d04-4f8e-b62f-69dd919bb3f4Cited by top-tier papers6
- Generative Video Motion Editing with 3D Point TracksYao-Chih Lee, Zhoutong Zhang, Jiahui Huang, Jui-Hsien Wang et al.CVPR 2026 · 23 citations
- BulletTime: Decoupled Control of Time and Camera Pose for Video GenerationYiming Wang, Qihang Zhang, Shengqu Cai, Tong Wu et al.CVPR 2026 · 15 citations
- WorldStereo: Bridging Camera-Guided Video Generation and Scene Reconstruction via 3D Geometric MemoriesYisu Zhang, Chenjie Cao, Tengfei Wang, Xuhui Zuo et al.CVPR 2026 · 13 citations
- LagerNVS: Latent Geometry for Fully Neural Real-time Novel View SynthesisStanislaw Szymanowicz, Minghao Chen, Jianyuan Wang, Christian Rupprecht et al.CVPR 2026 · 6 citations
- Guardians of the Hair: Rescuing Soft Boundaries in Depth, Stereo, and Novel ViewsXiang Zhang, Yang Zhang, Lukas Mehl, Markus Gross et al.CVPR 2026 · 2 citations
Builds on103
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov et al.ICCV 2023 · 1,662 citations
Related papers
- GSV3D: Gaussian Splatting-Based Geometric Distillation With Stable Video Diffusion for Single-Image 3D Object GenerationYe Tao, Jiawei Zhang, Yahao Shi, Dongqing Zou et al.ICCV 2025
- Generative Gaussian Splatting: Generating 3D Scenes with Video Diffusion PriorsKatja Schwarz, Norman Müller, Peter KontschiederICCV 2025 · 3 citations
- LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry GroundingJulian Ost, Andrea Ramazzina, Amogh Joshi, Maximilian Bömer et al.AAAI 2026 · 6 citations
- DistillNeRF: Perceiving 3D Scenes from Single-Glance Images by Distilling Neural Fields and Foundation Model FeaturesLetian Wang, Seung Wook Kim, Jiawei Yang, Cunjun Yu et al.NeurIPS 2024 · 35 citations
- GSFixer: Improving 3D Gaussian Splatting with Reference-Guided Video Diffusion PriorsXingyilang Yin, Qi Zhang, Jiahao Chang, Ying Feng et al.ICML 2026 · 33 citations
