ERUPT: Efficient Rendering with Unposed Patch Transformer
Maxim V. Shugaev, Vincent Chen, Maxim Karrenbach, Kyle Ashley, Bridget Kennedy, Naresh P. Cuntoor
摘要
This work addresses the problem of novel view synthesis in diverse scenes from small collections of RGB images. We propose ERUPT (Efficient Rendering with Unposed Patch Transformer) a state-of-the-art scene reconstruction model capable of efficient scene rendering using unposed imagery. We introduce patch-based querying, in contrast to existing pixel-based queries, to reduce the compute required to render a target view. This makes our model highly efficient both during training and at inference, capable of rendering at 600 fps on commercial hardware. Notably, our model is designed to use a learned latent camera pose which allows for training using unposed targets in datasets with sparse or inaccurate ground truth camera pose. We show that our approach can generalize on large real-world data and introduce a new benchmark dataset (MSVS-1M) for latent view synthesis using street-view imagery collected from Mapillary. In contrast to NeRF and Gaussian Splatting, which require dense imagery and precise metadata, ERUPT can render novel views of arbitrary scenes with as few as five unposed input images. ERUPT achieves better rendered image quality than current state-of-the-art methods for unposed image synthesis tasks, reduces labeled data requirements by 95% and decreases computational requirements by an order of magnitude, providing efficient novel view synthesis for diverse real-world scenes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 被引用 4,089 次
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 被引用 769 次
相关 Paper
- MuGS: Multi-Baseline Generalizable Gaussian Splatting ReconstructionYaopeng Lou, Li Shen, Tianqi Liu, Jiaqi Li 等ICCV 2025 · 被引用 1 次
- No Pose at All: Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse ViewsRanran Huang, Krystian MikolajczykICCV 2025 · 被引用 12 次
- SPARF: Neural Radiance Fields from Sparse and Noisy PosesPrune Truong, Marie-Julie Rakotosaona, Fabian Manhardt, Federico TombariCVPR 2023
- A Construct-Optimize Approach to Sparse View Synthesis without Camera PoseKaiwen Jiang, Yang Fu, Mukund Varma T., Yash Belhe 等SIGGRAPH 2024 · 被引用 20 次
- FewViewGS: Gaussian Splatting with Few View Matching and Multi-stage TrainingRuihong Yin, Vladimir Yugay, Yue Li, Sezer Karaoglu 等NeurIPS 2024 · 被引用 29 次
