True Self-Supervised Novel View Synthesis is Transferable
Thomas W. Mitchel, Hyunwoo Ryu, Vincent Sitzmann
Abstract
In this paper, we identify that the key criterion for determining whether a model is truly capable of novel view synthesis (NVS) is transferability: Whether any pose representation extracted from one video sequence can be used to re-render the same camera trajectory in another. We analyze prior work on self-supervised NVS and find that their predicted poses do not transfer: The same set of poses lead to different camera trajectories in different 3D scenes. Here, we present XFactor, the first geometry-free self-supervised model capable of true NVS. XFactor combines pair-wise pose estimation with a simple augmentation scheme of the inputs and outputs that jointly enables disentangling camera pose from scene content and facilitates geometric reasoning. Remarkably, we show that XFactor achieves transferability with unconstrained latent pose variables, without any 3D inductive biases or concepts from multi-view geometry -such as an explicit parameterization of poses as elements of SE(3). We introduce a new metric to quantify transferability, and through large-scale experiments, we demonstrate that XFactor significantly outperforms prior pose-free NVS transformers, and show that latent poses are highly correlated with real-world poses through probing experiments. Project
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 278f0102-4a51-4514-afd0-403fa0003e9dCited by top-tier papers6
- E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-trainingQitao Zhao, Hao Tan, Qianqian Wang, Sai Bi et al.CVPR 2026 · 24 citations
- SceneTok: A Compressed, Diffusable Token Space for 3D ScenesMohammad Asim, Christopher Wewer, Jan LenssenCVPR 2026 · 6 citations
- Scaling View Synthesis TransformersEvan Kim, Hyunwoo Ryu, Thomas W. Mitchel, Vincent SitzmannCVPR 2026 · 6 citations
- WildRayZer: Self-supervised Large View Synthesis in Dynamic EnvironmentsXuweiyi Chen, Wentao Zhou, Zezhou ChengCVPR 2026 · 5 citations
- From None to All: Self-Supervised 3D Reconstruction via Novel View SynthesisRanran Huang, Weixun Luo, Ye Mao, Krystian MikolajczykCVPR 2026 · 2 citations
Builds on23
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised LearningAdrien Bardes, Jean Ponce, Yann LeCunICLR 2022 · 1,226 citations
- Common Objects in 3D: Large-Scale Learning and Evaluation of Real-life 3D Category ReconstructionJeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone et al.ICCV 2021 · 686 citations
- Light Field Networks: Neural Scene Representations with Single-Evaluation RenderingVincent Sitzmann, Semon Rezchikov, Bill Freeman, Josh Tenenbaum et al.NeurIPS 2021 · 426 citations
- CroCo v2: Improved Cross-view Completion Pre-training for Stereo Matching and Optical FlowPhilippe Weinzaepfel, Thomas Lucas, Vincent Leroy, Yohann Cabon et al.ICCV 2023 · 181 citations
Related papers
- Rayzer: a Self-Supervised Large View Synthesis ModelHanwen Jiang, Hao Tan, Peng Wang, Hai Jin et al.ICCV 2025 · 12 citations
- Video Autoencoder: self-supervised disentanglement of static 3D structure and motionZihang Lai, Sifei Liu, Alexei A. Efros, Xiaolong WangICCV 2021 · 37 citations
- MAPConNet: Self-supervised 3D Pose Transfer with Mesh and Point Contrastive LearningJiaze Sun, Zhixiang Chen, Tae-Kyun KimICCV 2023 · 2 citations
- RUST: Latent Neural Scene Representations from Unposed ImageryMehdi S. M. Sajjadi, Aravindh Mahendran, Thomas Kipf, Etienne Pot et al.CVPR 2023
- Epipolar-Free 3D Gaussian Splatting for Generalizable Novel View SynthesisZhiyuan Min, Yawei Luo, Jianwen Sun, Yi YangNeurIPS 2024 · 23 citations
