LVSM: A Large View Synthesis Model with Minimal 3D Inductive Bias
Haian Jin, Hanwen Jiang, Hao Tan, Kai Zhang, Sai Bi, Tianyuan Zhang, Fujun Luan, Noah Snavely, Zexiang Xu
Abstract
We propose the Large View Synthesis Model (LVSM), a novel transformer-based approach for scalable and generalizable novel view synthesis from sparse-view inputs. We introduce two architectures: (1) an encoder-decoder LVSM, which encodes input image tokens into a fixed number of 1D latent tokens, functioning as a fully learned scene representation, and decodes novel-view images from them; and (2) a decoder-only LVSM, which directly maps input images to novelview outputs, completely eliminating intermediate scene representations. Both models bypass the 3D inductive biases used in previous methods-from 3D representations (e.g., NeRF, 3DGS) to network designs (e.g., epipolar projections, plane sweeps)-addressing novel view synthesis with a fully data-driven approach. While the encoder-decoder model offers faster inference due to its independent latent representation, the decoder-only LVSM achieves superior quality, scalability, and zero-shot generalization, outperforming previous state-of-the-art methods by 1.5 to 3.5 dB PSNR. Comprehensive evaluations across multiple datasets demonstrate that both LVSM variants achieve state-of-the-art novel view synthesis quality. Notably, our models surpass all previous methods even with reduced computational resources (1-2 GPUs). Please see our website for more results: https://haian-jin.github.io/projects/LVSM/ . RELATED WORK View Synthesis. Novel view synthesis (NVS) has been studied for decades. Image-based rendering (IBR) methods perform view synthesis by weighted blending of input reference images using proxy
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8b9f225d-b13d-41b6-b200-c39cf9efd377Cited by top-tier papers67
- Test-Time Training Done RightTianyuan Zhang, Sai Bi, Yicong Hong, Kai Zhang et al.ICLR 2026 · 127 citations
- MVSplat360: Feed-Forward 360 Scene Synthesis from Sparse ViewsYuedong Chen, Chuanxia Zheng, Haofei Xu, Bohan Zhuang et al.NeurIPS 2024 · 126 citations
- Cameras as Relative Positional EncodingRuilong Li, Brent Yi, Junchen Liu, Hang Gao et al.NeurIPS 2025 · 113 citations
- Feed-Forward Bullet-Time Reconstruction of Dynamic Scenes from Monocular VideosHanxue Liang, Jiawei Ren, Ashkan Mirzaei, Antonio Torralba et al.NeurIPS 2025 · 52 citations
- Dens3R: A Foundation Model for 3D Geometry PredictionXianze Fang, Jingnan Gao, Zhe Wang, Zhuo Chen et al.ICLR 2026 · 45 citations
Builds on48
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
Related papers
- Efficient-LVSM: Faster, Cheaper, and Better Large View Synthesis Model via Decoupled Co-Refinement AttentionXiaosong Jia, Yihang Sun, Junqi You, Songbur Wong et al.ICLR 2026 · 6 citations
- Scaling Transformer-Based Novel View Synthesis with Models Token Disentanglement and Synthetic DataNithin Gopalakrishnan Nair, Srinivas Kaza, Xuan Luo, Vishal M. Patel et al.ICCV 2025 · 1 citation
- Scaling View Synthesis TransformersEvan Kim, Hyunwoo Ryu, Thomas W. Mitchel, Vincent SitzmannCVPR 2026 · 6 citations
- UMAMI: Unifying Masked Autoregressive Models and Deterministic Rendering for View SynthesisThanh-Tung Le, Tuan Pham, Tung Nguyen, Deying Kong et al.NeurIPS 2025 · 4 citations
- SparseGNV: Generating Novel Views of Indoor Scenes with Sparse RGB-D ImagesWeihao Cheng, Yan-Pei Cao, Ying ShanAAAI 2024 · 2 citations
