NViST: In the Wild New View Synthesis from a Single Image with Transformers
Wonbong Jang, Lourdes Agapito
Abstract
We propose NViST, a transformer-based model for efficient and generalizable novel-view synthesis from a single image for real-world scenes. In contrast to many methods that are trained on synthetic data, object-centred scenarios, or in a category-specific manner, NViST is trained on MVImgNet, a large-scale dataset of casually-captured real-world videos of hundreds of object categories with diverse backgrounds. NViST transforms image inputs directly into a radiance field, conditioned on camera parameters via adaptive layer normalisation. In practice, NViST exploits fine-tuned masked autoencoder (MAE) features and translates them to 3D output tokens via cross-attention, while addressing occlusions with self-attention. To move away from object-centred datasets and enable full scene synthesis, NViST adopts a 6-DOF camera pose model and only requires relative pose, dropping the need for canonicalization of the training data, which removes a substantial barrier to it being used on casually captured datasets. We show results on unseen objects and categories from MVImgNet and even generalization to casual phone captures. We conduct qualitative and quantitative evaluations on MVImgNet and ShapeNet to show that our model represents a step forward towards enabling true in-the-wild generalizable novel-view synthesis from a single image. Project webpage: https://wbjang.github.io/nvist_webpage.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3e2b4da-377d-49de-b7b4-6dd4331932a6Cited by top-tier papers6
- Scaling Sequence-to-Sequence Generative Neural RenderingShikun Liu, Kam Woh Ng, Wonbong Jang, Jiadong Guo et al.ICLR 2026 · 9 citations
- Rays as Pixels: Learning A Joint Distribution of Video and Camera TrajectoriesWonbong Jang, Shikun Liu, Soubhik Sanyal, Juan Perez et al.ICML 2026 · 3 citations
- Novel View Synthesis with Pixel-Space Diffusion ModelsNoam Elata, Bahjat Kawar, Yaron Ostrovsky-Berman, Miriam Farber et al.CVPR 2025
- SpatialDreamer: Self-supervised Stereo Video Synthesis from Monocular InputZhen Lv, Yangqi Long, Congzhentao Huang, Cao Li et al.CVPR 2025
- Inverse Image-Based Rendering for Light Field Generation From Single ImagesHyunjun Jung, Hae-Gon JeonICCV 2025
Builds on66
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- Free3D: Consistent Novel View Synthesis Without 3D RepresentationChuanxia Zheng, Andrea VedaldiCVPR 2024 · 28 citations
- LVSM: A Large View Synthesis Model with Minimal 3D Inductive BiasHaian Jin, Hanwen Jiang, Hao Tan, Kai Zhang et al.ICLR 2025
- RUST: Latent Neural Scene Representations from Unposed ImageryMehdi S. M. Sajjadi, Aravindh Mahendran, Thomas Kipf, Etienne Pot et al.CVPR 2023
- MVInpainter: Learning Multi-View Consistent Inpainting to Bridge 2D and 3D EditingChenjie Cao, Chaohui Yu, Fan Wang, Xiangyang Xue et al.NeurIPS 2024 · 36 citations
- MultiDiff: Consistent Novel View Synthesis from a Single ImageNorman Müller, Katja Schwarz, Barbara Rössle, Lorenzo Porzi et al.CVPR 2024 · 14 citations
