PlaneTR: Structure-Guided Transformers for 3D Plane Recovery
Bin Tan, Nan Xue, Song Bai, Tianfu Wu, Gui-Song Xia
摘要
This paper presents a neural network built upon Transformers, namely PlaneTR, to simultaneously detect and reconstruct planes from a single image. Different from previous methods, PlaneTR jointly leverages the context information and the geometric structures in a sequence-to-sequence way to holistically detect plane instances in one forward pass. Specifically, we represent the geometric structures as line segments and conduct the network with three main components: (i) context and line segments encoders, (ii) a structure-guided plane decoder, (iii) a pixel-wise plane embedding decoder. Given an image and its detected line segments, PlaneTR generates the context and line segment sequences via two specially designed encoders and then feeds them into a Transformers-based decoder to directly predict a sequence of plane instances by simultaneously considering the context and global structure cues. Finally, the pixel-wise embeddings are computed to assign each pixel to one predicted plane instance which is nearest to it in embedding space. Comprehensive experiments demonstrate that PlaneTR achieves state-of-the-art performance on the ScanNet and NYUv2 datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- MonoDTR: Monocular 3D Object Detection with Depth-Aware TransformerKuan-Chih Huang, Tsung-Han Wu, Hung-Ting Su, Winston H. HsuCVPR 2022 · 被引用 199 次
- Unsupervised Homography Estimation with Coplanarity-Aware GANMingbo Hong, Yuhang Lu, Nianjin Ye, Chunyu Lin 等CVPR 2022 · 被引用 62 次
- PlaneMVS: 3D Plane Reconstruction from Multi-View StereoJiachen Liu, Pan Ji, Nitin Bansal, Changjiang Cai 等CVPR 2022 · 被引用 43 次
- PlanarRecon: Realtime 3D Plane Detection and Reconstruction from Posed Monocular VideosYiming Xie, Matheus Gadelha, Fengting Yang, Xiaowei Zhou 等CVPR 2022 · 被引用 32 次
- ExtDM: Distribution Extrapolation Diffusion Model for Video PredictionZhicheng Zhang, Junyao Hu, Wentao Cheng, Danda Pani Paudel 等CVPR 2024 · 被引用 24 次
它引用的顶会 Paper12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- End-to-End Wireframe ParsingYichao Zhou, Haozhi Qi, Yi MaICCV 2019 · 被引用 190 次
- Multi-Plane Program Induction with 3D Box PriorsYikai Li, Jiayuan Mao, Xiuming Zhang, Bill Freeman 等NeurIPS 2020 · 被引用 18 次
- Geometric Structure Based and Regularized Depth Estimation From 360 Indoor ImageryLei Jin, Yanyu Xu, Jia Zheng, Junfei Zhang 等CVPR 2020
相关 Paper
- PlaneRecTR: Unified Query Learning for 3D Plane Recovery from a Single ViewJingjia Shi, Shuaifeng Zhi, Kai XuICCV 2023
- Towards In-the-wild 3D Plane Reconstruction from a Single ImageJiachen Liu, Rui Yu, Sili Chen, Sharon X. Huang 等CVPR 2025
- Uni-3D: A Universal Model for Panoptic 3D Scene ReconstructionXiang Zhang, Zeyuan Chen, Fangyin Wei, Zhuowen TuICCV 2023 · 被引用 24 次
- Connecting the Dots: Floorplan Reconstruction Using Two-Level QueriesYuanwen Yue, Theodora Kontogianni, Konrad Schindler, Francis EngelmannCVPR 2023
- PLANA3R: Zero-shot Metric Planar 3D Reconstruction via Feed-forward Planar SplattingChangkun Liu, Bin Tan, Zeran Ke, Shangzhan Zhang 等NeurIPS 2025 · 被引用 6 次
