RUST: Latent Neural Scene Representations from Unposed Imagery
Mehdi S. M. Sajjadi, Aravindh Mahendran, Thomas Kipf, Etienne Pot, Daniel Duckworth, Mario Lucic, Klaus Greff
Abstract
Inferring the structure of 3D scenes from 2D observations is a fundamental challenge in computer vision. Recently popularized approaches based on neural scene representations have achieved tremendous impact and have been applied across a variety of applications. One of the major remaining challenges in this space is training a single model which can provide latent representations which effectively generalize beyond a single scene. Scene Representation Transformer (SRT) has shown promise in this direction, but scaling it to a larger set of diverse scenes is challenging and necessitates accurately posed ground truth data. To address this problem, we propose RUST (Really Unposed Scene representation Transformer), a pose-free approach to novel view synthesis trained on RGB images alone. Our main insight is that one can train a Pose Encoder that peeks at the target image and learns a latent pose embedding which is used by the decoder for view synthesis. We perform an empirical investigation into the learned latent pose structure and show that it allows meaningful test-time camera transformations and accurate explicit pose readouts. Perhaps surprisingly, RUST achieves similar quality as methods which have access to perfect camera pose, thereby unlocking the potential for large-scale training of amortized neural scene representations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers24
- Large Spatial Model: End-to-end Unposed Images to Semantic 3DZhiwen Fan, Jian Zhang, Wenyan Cong, Peihao Wang et al.NeurIPS 2024 · 86 citations
- Efficiently Reconstructing Dynamic Scenes One D4RT at a TimeChuhan Zhang, Guillaume Le Moing, Skanda Koppula, Ignacio Rocco et al.CVPR 2026 · 52 citations
- LU-NeRF: Scene and Pose Estimation by Synchronizing Local Unposed NeRFsZezhou Cheng, Carlos Esteves, Varun Jampani, Abhishek Kar et al.ICCV 2023 · 46 citations
- FlowCam: Training Generalizable 3D Radiance Fields without Camera Poses via Pixel-Aligned Scene FlowCameron Smith, Yilun Du, Ayush Tewari, Vincent SitzmannNeurIPS 2023 · 43 citations
- E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-trainingQitao Zhao, Hao Tan, Qianqian Wang, Sai Bi et al.CVPR 2026 · 24 citations
Builds on12
- BARF: Bundle-Adjusting Neural Radiance FieldsChen-Hsuan Lin, Wei-Chiu Ma, Antonio Torralba, Simon LuceyICCV 2021 · 867 citations
- Putting NeRF on a Diet: Semantically Consistent Few-Shot View SynthesisAjay Jain, Matthew Tancik, Pieter AbbeelICCV 2021 · 615 citations
- RegNeRF: Regularizing Neural Radiance Fields for View Synthesis from Sparse InputsMichael Niemeyer, Jonathan T. Barron, Ben Mildenhall, Mehdi S. M. Sajjadi et al.CVPR 2022 · 513 citations
- GNeRF: GAN-based Neural Radiance Field without Posed CameraQuan Meng, Anpei Chen, Haimin Luo, Minye Wu et al.ICCV 2021 · 222 citations
- Kubric: A scalable dataset generatorKlaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch et al.CVPR 2022 · 183 citations
Related papers
- Scene Representation Transformer: Geometry-Free Novel View Synthesis Through Set-Latent Scene RepresentationsMehdi S. M. Sajjadi, Henning Meyer, Etienne Pot, Urs Bergmann et al.CVPR 2022 · 102 citations
- AutoRF: Learning 3D Object Radiance Fields from Single View ObservationsNorman Müller, Andrea Simonelli, Lorenzo Porzi, Samuel Rota Bulò et al.CVPR 2022 · 45 citations
- Equivariant Neural RenderingEmilien Dupont, Miguel Bautista Martin, Alex Colburn, Aditya Sankar et al.ICML 2020 · 69 citations
- Rayzer: a Self-Supervised Large View Synthesis ModelHanwen Jiang, Hao Tan, Peng Wang, Hai Jin et al.ICCV 2025 · 12 citations
- NViST: In the Wild New View Synthesis from a Single Image with TransformersWonbong Jang, Lourdes AgapitoCVPR 2024
