Towards In-the-wild 3D Plane Reconstruction from a Single Image
Jiachen Liu, Rui Yu, Sili Chen, Sharon X. Huang, Hengkai Guo
Abstract
3D plane reconstruction from a single image is a crucial yet challenging topic in 3D computer vision. Previous stateof-the-art (SOTA) methods have focused on training their system on a single dataset from either indoor or outdoor domain, limiting their generalizability across diverse testing data. In this work, we introduce a novel framework dubbed ZeroPlane, a Transformer-based model targeting zero-shot 3D plane detection and reconstruction from a single image, over diverse domains and environments. To enable datadriven models across multiple domains, we have curated a large-scale planar benchmark, comprising over 14 datasets and 560,000 high-resolution, dense planar annotations for diverse indoor and outdoor scenes. To address the challenge of achieving desirable planar geometry on multi-dataset training, we propose to disentangle the representation of plane normal and offset, and employ an exemplar-guided, classification-then-regression paradigm to learn plane and offset respectively. Additionally, we employ advanced backbones as image encoder, and present an effective pixelgeometry-enhanced plane embedding module to further facilitate planar reconstruction. Extensive experiments across multiple zero-shot evaluation datasets have demonstrated that our approach significantly outperforms previous methods on both reconstruction accuracy and generalizability, especially over in-the-wild data. Our code and data are available at: https://github.com/jcliu0428/ZeroPlane .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a850ffab-d6e7-486b-bcb0-07273c373e75Cited by top-tier papers5
- G4Splat: Geometry-Guided Gaussian Splatting with Generative PriorJunfeng Ni, Yixin Chen, Zhifei Yang, Yu Liu et al.ICLR 2026 · 10 citations
- PLANA3R: Zero-shot Metric Planar 3D Reconstruction via Feed-forward Planar SplattingChangkun Liu, Bin Tan, Zeran Ke, Shangzhan Zhang et al.NeurIPS 2025 · 6 citations
- PlanarGS: High-Fidelity Indoor 3D Gaussian Splatting Guided by Vision-Language Planar PriorsXirui Jin, Renbiao Jin, Boying Li, Danping Zou et al.NeurIPS 2025 · 5 citations
- PlanaReLoc: Camera Relocalization in 3D Planar Primitives via Region-Based Structure MatchingHanqiao Ye, Yuzhou Liu, Yangdong Liu, Shuhan ShenCVPR 2026
- Top2Pano: Learning to Generate Indoor Panoramas from Top-Down ViewZitong Zhang, Suranjan Gautam, Rui YuICCV 2025
Builds on26
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
Related papers
- PlaneTR: Structure-Guided Transformers for 3D Plane RecoveryBin Tan, Nan Xue, Song Bai, Tianfu Wu et al.ICCV 2021 · 51 citations
- PlaneRecTR: Unified Query Learning for 3D Plane Recovery from a Single ViewJingjia Shi, Shuaifeng Zhi, Kai XuICCV 2023
- PlaneMVS: 3D Plane Reconstruction from Multi-View StereoJiachen Liu, Pan Ji, Nitin Bansal, Changjiang Cai et al.CVPR 2022 · 43 citations
- NeuralPlane: Structured 3D Reconstruction in Planar Primitives with Neural FieldsHanqiao Ye, Yuzhou Liu, Yangdong Liu, Shuhan ShenICLR 2025
- ZeroShape: Regression-Based Zero-Shot Shape ReconstructionZixuan Huang, Stefan Stojanov, Anh Thai, Varun Jampani et al.CVPR 2024
