LGT-Net: Indoor Panoramic Room Layout Estimation with Geometry-Aware Transformer Network
Zhigang Jiang, Zhongzheng Xiang, Jinhua Xu, Ming Zhao
Abstract
3D room layout estimation by a single panorama using deep neural networks has made great progress. However, previous approaches can not obtain efficient geometry awareness of room layout with the only latitude of boundaries or horizon-depth. We present that using horizondepth along with room height can obtain omnidirectionalgeometry awareness of room layout in both horizontal and vertical directions. In addition, we propose a planar-geometry aware loss function with normals and gradients of normals to supervise the planeness of walls and turning of corners. We propose an efficient network, LGT-Net, for room layout estimation, which contains a novel Transformer architecture called SWG-Transformer to model geometry relations. SWG-Transformer consists of (Shifted) Window Blocks and Global Blocks to combine the local and global geometry relations. Moreover, we design a novel relative position embedding of Transformer to enhance the spatial identification ability for the panorama. Experiments show that the proposed LGT-Net achieves better performance than current state-of-the-arts (SOTA) on benchmark datasets. The code is publicly available at https://github.com/zhigangjiang/LGT-Net.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b35a34c-7987-4cda-aa8c-0d588d4eb7f2Cited by top-tier papers9
- Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and GenerationKang Liao, Size Wu, Zhonghua Wu, Linyi Jin et al.ICLR 2026 · 19 citations
- CAGE: Continuity-Aware edGE Network Unlocks Robust Floorplan ReconstructionYiyi Liu, Chunyang Liu, Bohan Wang, Weiqin Jiao et al.NeurIPS 2025 · 7 citations
- No More Ambiguity in 360° Room Layout via Bi-Layout EstimationYu-Ju Tsai, Jin-Cheng Jhang, Jingjing Zheng, Wei Wang et al.CVPR 2024 · 6 citations
- PanoPose: Self-supervised Relative Pose Estimation for Panoramic ImagesDiantao Tu, Hainan Cui, Xianwei Zheng, Shuhan ShenCVPR 2024 · 4 citations
- PanelNet: Understanding 360 Indoor Environment via Panel RepresentationHaozheng Yu, Lu He, Bing Jian, Weiwei Feng et al.CVPR 2023
Builds on6
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
- Rethinking Positional Encoding in Language Pre-trainingGuolin Ke, Di He, Tie-Yan LiuICLR 2021 · 358 citations
Related papers
- PSMNet: Position-aware Stereo Merging Network for Room Layout EstimationHaiyan Wang, Will Hutchcroft, Yuguang Li, Zhiqiang Wan et al.CVPR 2022 · 20 citations
- LED2-Net: Monocular 360deg Layout Estimation via Differentiable Depth RenderingFu-En Wang, Yu-Hsuan Yeh, Min Sun, Wei-Chen Chiu et al.CVPR 2021
- HoHoNet: 360 Indoor Holistic Understanding With Latent Horizontal FeaturesCheng Sun, Min Sun, Hwann-Tzong ChenCVPR 2021
- PanoContext-Former: Panoramic Total Scene Understanding with a TransformerYuan Dong, Chuan Fang, Liefeng Bo, Zilong Dong et al.CVPR 2024
- PanoSwin: a Pano-style Swin Transformer for Panorama UnderstandingZhixin Ling, Zhen Xing, Xiangdong Zhou, Manliang Cao et al.CVPR 2023
