Worldsheet: Wrapping the World in a 3D Sheet for View Synthesis from a Single Image
Ronghang Hu, Nikhila Ravi, Alexander C. Berg, Deepak Pathak
Abstract
We present Worldsheet, a method for novel view synthesis using just a single RGB image as input. The main insight is that simply shrink-wrapping a planar mesh sheet onto the input image, consistent with the learned intermediate depth, captures underlying geometry sufficient to generate photorealistic unseen views with large viewpoint changes. To operationalize this, we propose a novel differentiable texture sampler that allows our wrapped mesh sheet to be textured and rendered differentiably into an image from a target viewpoint. Our approach is category-agnostic, end-to-end trainable without using any 3D supervision, and requires a single image at test time. We also explore a simple extension by stacking multiple layers of Worldsheets to better handle occlusions. Worldsheet consistently outperforms prior state-of-the-art methods on single-image view synthesis across several datasets. Furthermore, this simple idea captures novel views surprisingly well on a wide range of high-resolution in-the-wild images, converting them into navigable 3D pop-ups. Video results and code are available at https://worldsheet.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 42780652-2b0a-4390-8b73-c2d86d5b9113Cited by top-tier papers44
- BARF: Bundle-Adjusting Neural Radiance FieldsChen-Hsuan Lin, Wei-Chiu Ma, Antonio Torralba, Simon LuceyICCV 2021 · 867 citations
- Putting NeRF on a Diet: Semantically Consistent Few-Shot View SynthesisAjay Jain, Matthew Tancik, Pieter AbbeelICCV 2021 · 615 citations
- SceneScape: Text-Driven Consistent Scene GenerationRafail Fridman, Amit Abecasis, Yoni Kasten, Tali DekelNeurIPS 2023 · 196 citations
- Geometry-Free View Synthesis: Transformers and no 3D PriorsRobin Rombach, Patrick Esser, Björn OmmerICCV 2021 · 115 citations
- PixelSynth: Generating a 3D-Consistent Experience from a Single ImageChris Rockwell, David F. Fouhey, Justin JohnsonICCV 2021 · 98 citations
Builds on15
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- Neural Sparse Voxel FieldsLingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua et al.NeurIPS 2020 · 1,535 citations
- Soft Rasterizer: A Differentiable Renderer for Image-Based 3D ReasoningShichen Liu, Weikai Chen, Tianye Li, Hao LiICCV 2019 · 789 citations
- TexturePose: Supervising Human Mesh Estimation With Texture ConsistencyGeorgios Pavlakos, Nikos Kolotouros, Kostas DaniilidisICCV 2019 · 109 citations
- Canonical Surface Mapping via Geometric Cycle ConsistencyNilesh Kulkarni, Shubham Tulsiani, Abhinav GuptaICCV 2019 · 104 citations
Related papers
- SynSin: End-to-End View Synthesis From a Single ImageOlivia Wiles, Georgia Gkioxari, Richard Szeliski, Justin JohnsonCVPR 2020
- Novel View Synthesis with Pixel-Space Diffusion ModelsNoam Elata, Bahjat Kawar, Yaron Ostrovsky-Berman, Miriam Farber et al.CVPR 2025
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov et al.ICCV 2023 · 1,662 citations
- TMO: Textured Mesh Acquisition of Objects with a Mobile Device by using Differentiable RenderingJaehoon Choi, Dongki Jung, Taejae Lee, Sangwook Kim et al.CVPR 2023
- Differentiable Volumetric Rendering: Learning Implicit 3D Representations Without 3D SupervisionMichael Niemeyer, Lars M. Mescheder, Michael Oechsle, Andreas GeigerCVPR 2020
