Generating 3D-Consistent Videos from Unposed Internet Photos
Gene Chou, Kai Zhang, Sai Bi, Hao Tan, Zexiang Xu, Fujun Luan, Bharath Hariharan, Noah Snavely
Abstract
Figure 1. Given n unposed input keyframes, the goal is to generate a video of the scene with a realistic camera trajectory and consistent geometry. From top to bottom: Ours, Luma Dream Machine [41] (a commercial video generation model), FILM [50] (a frame interpolation method). Luma hallucinates new buildings (left scene) and statues (right scene) without understanding the scene layout. FILM is unable to handle wide-baseline inputs and produces blurry transitions. See our supplement for video playback.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80cbc0e1-c103-4615-a171-6c3cdf01c32cCited by top-tier papers2
- Reangle-A-Video: 4D Video Generation as Video-to-Video TranslationHyeonho Jeong, Suhyeon Lee, Jong Chul YeICCV 2025 · 3 citations
- FlashDepth: Real-Time Streaming Video Depth Estimation at 2K ResolutionGene Chou, Wenqi Xian, Guandao Yang, Mohamed Abdelfattah et al.ICCV 2025 · 1 citation
Builds on47
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
Related papers
- ViBiDSampler: Enhancing Video Interpolation Using Bidirectional Diffusion SamplerSerin Yang, Taesung Kwon, Jong Chul YeICLR 2025
- Generative Inbetweening through Frame-wise Conditions-Driven Video GenerationTianyi Zhu, Dongwei Ren, Qilong Wang, Xiaohe Wu et al.CVPR 2025
- Geometry-as-context: Modulating Explicit 3D in Scene-consistent Video Generation to Geometry ContextJiaKui Hu, Jialun Liu, Liying Yang, Xinliang Zhang et al.CVPR 2026 · 7 citations
- SketchVideo: Sketch-based Video Generation and EditingFeng-Lin Liu, Hongbo Fu, Xintao Wang, Weicai Ye et al.CVPR 2025
- Consistent View Synthesis with Pose-Guided Diffusion ModelsHung-Yu Tseng, Qinbo Li, Changil Kim, Suhib Alsisan et al.CVPR 2023
