Geometry-Free View Synthesis: Transformers and no 3D Priors
Robin Rombach, Patrick Esser, Björn Ommer
摘要
Is a geometric model required to synthesize novel views from a single image? Being bound to local convolutions, CNNs need explicit 3D biases to model geometric transformations. In contrast, we demonstrate that a transformer-based model can synthesize entirely novel views without any hand-engineered 3D biases. This is achieved by (i) a global attention mechanism for implicitly learning long-range 3D correspondences between source and target views, and (ii) a probabilistic formulation necessary to capture the ambiguity inherent in predicting novel views from a single image, thereby overcoming the limitations of previous approaches that are restricted to relatively small viewpoint changes. We evaluate various ways to integrate 3D priors into a transformer architecture. However, our experiments show that no such geometric priors are required and that the transformer is capable of implicitly learning 3D relationships between images. Furthermore, this approach outperforms the state of the art in terms of visual quality while covering the full distribution of possible realizations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper51
- Generative Novel View Synthesis with 3D-Aware Diffusion ModelsEric R. Chan, Koki Nagano, Matthew A. Chan, Alexander W. Bergman 等ICCV 2023 · 被引用 314 次
- NerfDiff: Single-image View Synthesis with NeRF-guided Distillation from 3D-aware DiffusionJiatao Gu, Alex Trevithick, Kai-En Lin, Joshua M. Susskind 等ICML 2023 · 被引用 224 次
- SceneScape: Text-Driven Consistent Scene GenerationRafail Fridman, Amit Abecasis, Yoni Kasten, Tali DekelNeurIPS 2023 · 被引用 196 次
- GAUDI: A Neural Architect for Immersive 3D Scene GenerationMiguel Ángel Bautista, Pengsheng Guo, Samira Abnar, Walter Talbott 等NeurIPS 2022 · 被引用 170 次
- Object Scene Representation TransformerMehdi S. M. Sajjadi, Daniel Duckworth, Aravindh Mahendran, Sjoerd van Steenkiste 等NeurIPS 2022 · 被引用 124 次
它引用的顶会 Paper15
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 被引用 2,416 次
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu 等ICML 2020 · 被引用 1,773 次
- Infinite Nature: Perpetual View Generation of Natural Scenes from a Single ImageAndrew Liu, Ameesh Makadia, Richard Tucker, Noah Snavely 等ICCV 2021 · 被引用 260 次
- Extreme View SynthesisInchang Choi, Orazio Gallo, Alejandro J. Troccoli, Min H. Kim 等ICCV 2019 · 被引用 207 次
相关 Paper
- GTA: A Geometry-Aware Attention Mechanism for Multi-View TransformersTakeru Miyato, Bernhard Jaeger, Max Welling, Andreas GeigerICLR 2024 · 被引用 51 次
- Look Outside the Room: Synthesizing A Consistent Long-Term 3D Scene Video from A Single ImageXuanchi Ren, Xiaolong WangCVPR 2022 · 被引用 42 次
- Rayzer: a Self-Supervised Large View Synthesis ModelHanwen Jiang, Hao Tan, Peng Wang, Hai Jin 等ICCV 2025 · 被引用 12 次
- LVSM: A Large View Synthesis Model with Minimal 3D Inductive BiasHaian Jin, Hanwen Jiang, Hao Tan, Kai Zhang 等ICLR 2025
- Is Attention All That NeRF Needs?Mukund Varma T., Peihao Wang, Xuxi Chen, Tianlong Chen 等ICLR 2023 · 被引用 6 次
