Rayzer: a Self-Supervised Large View Synthesis Model
Hanwen Jiang, Hao Tan, Peng Wang, Hai Jin, Yue Zhao, Sai Bi, Kai Zhang, Fujun Luan, Kalyan Sunkavalli, Qixing Huang, Georgios Pavlakos
Abstract
We present RayZer, a self-supervised multi-view 3D Vision model trained without any 3D supervision, i.e., camera poses and scene geometry, while exhibiting emerging 3D awareness. Concretely, RayZer takes unposed and uncalibrated images as input, recovers camera parameters, reconstructs a scene representation, and synthesizes novel views. During training, RayZer relies solely on its self-predicted camera poses to render target views, eliminating the need for any ground-truth camera annotations and allowing RayZer to be trained with 2D image supervision. The emerging 3D awareness of RayZer is attributed to two key factors. First, we design a self-supervised framework, which achieves 3D-aware auto-encoding of input images by disentangling camera and scene representations. Second, we design a transformerbased model in which the only 3D prior is the ray structure, connecting camera, pixel, and scene simultaneously. RayZer demonstrates comparable or even superior novel view synthesis performance than “oracle” methods that rely on pose annotations in both training and testing. Project: https://hwjiang1510.github.io/RayZer/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 47d4f82f-191b-406f-9413-2de8ea0af4e8Cited by top-tier papers32
- WorldMem: Long-term Consistent World Simulation with MemoryZeqi Xiao, Yushi Lan, Yifan Zhou, Wenqi Ouyang et al.NeurIPS 2025 · 165 citations
- UFM: A Simple Path towards Unified Dense Correspondence with FlowYuchen Zhang, Nikhil Varma Keetha, Chenwei Lyu, Bhuvan Jhamb et al.NeurIPS 2025 · 40 citations
- Unified Camera Positional Encoding for Controlled Video GenerationCheng Zhang, Boying Li, Meng Wei, Yan-Pei Cao et al.CVPR 2026 · 38 citations
- E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-trainingQitao Zhao, Hao Tan, Qianqian Wang, Sai Bi et al.CVPR 2026 · 24 citations
- ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time TrainingHaian Jin, Rundi Wu, Tianyuan Zhang, Ruiqi Gao et al.CVPR 2026 · 23 citations
Builds on47
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
Related papers
- From None to All: Self-Supervised 3D Reconstruction via Novel View SynthesisRanran Huang, Weixun Luo, Ye Mao, Krystian MikolajczykCVPR 2026 · 2 citations
- WildRayZer: Self-supervised Large View Synthesis in Dynamic EnvironmentsXuweiyi Chen, Wentao Zhou, Zezhou ChengCVPR 2026 · 5 citations
- ViewNet: Unsupervised Viewpoint Estimation from Conditional GenerationOctave Mariotti, Oisin Mac Aodha, Hakan BilenICCV 2021 · 8 citations
- Video Autoencoder: self-supervised disentanglement of static 3D structure and motionZihang Lai, Sifei Liu, Alexei A. Efros, Xiaolong WangICCV 2021 · 37 citations
- True Self-Supervised Novel View Synthesis is TransferableThomas W. Mitchel, Hyunwoo Ryu, Vincent SitzmannICLR 2026 · 13 citations
