Scaling Transformer-Based Novel View Synthesis with Models Token Disentanglement and Synthetic Data
Nithin Gopalakrishnan Nair, Srinivas Kaza, Xuan Luo, Vishal M. Patel, Stephen Lombardi, Jungyeon Park
摘要
Large transformer-based models have made significant progress in generalizable novel view synthesis (NVS) from sparse input views, generating novel viewpoints without the need for test-time optimization. However, these models are constrained by the limited diversity of publicly available scene datasets, making most real-world (in-the-wild) scenes out-of-distribution. To overcome this, we incorporate synthetic training data generated from diffusion models, which improves generalization across unseen domains. While synthetic data offers scalability, we identify artifacts introduced during data generation as a key bottleneck affecting reconstruction quality. To address this, we propose a token disentanglement process within the transformer architecture, enhancing feature separation and ensuring more effective learning. This refinement not only improves reconstruction quality over standard transformers but also enables scalable training with synthetic data. As a result, our method outperforms existing models on both in-dataset and cross-dataset evaluations, achieving state-of-the-art results across multiple benchmarks while significantly reducing computational costs. Project page: https://scaling3dnvs.github.io/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Multi-view Pyramid Transformer: Look Coarser to See BroaderGyeongjin Kang, Seungkwon Yang, Seungtae Nam, Younggeun Lee 等CVPR 2026 · 被引用 8 次
- FreeScale: Scaling 3D Scenes via Certainty-Aware Free-View GenerationChenhan Jiang, Yu Chen, Qingwen Zhang, Jifei Song 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper34
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- UMAMI: Unifying Masked Autoregressive Models and Deterministic Rendering for View SynthesisThanh-Tung Le, Tuan Pham, Tung Nguyen, Deying Kong 等NeurIPS 2025 · 被引用 4 次
- LVSM: A Large View Synthesis Model with Minimal 3D Inductive BiasHaian Jin, Hanwen Jiang, Hao Tan, Kai Zhang 等ICLR 2025
- HAD: Hallucination-Aware Diffusion Priors for 3D ReconstructionXi Liu, Weiwei Sun, Zhou Ren, Chris Broaddus 等CVPR 2026 · 被引用 1 次
- SemanticNVS: Improving Semantic Scene Understanding in Generative Novel View SynthesisXinya Chen, Christopher Wewer, Jiahao Xie, Xinting Hu 等ICML 2026
- SceneTok: A Compressed, Diffusable Token Space for 3D ScenesMohammad Asim, Christopher Wewer, Jan LenssenCVPR 2026 · 被引用 6 次
