PanoContext-Former: Panoramic Total Scene Understanding with a Transformer
Yuan Dong, Chuan Fang, Liefeng Bo, Zilong Dong, Ping Tan
摘要
Panoramic images enable deeper understanding and more holistic perception of 360 • surrounding environment, which can naturally encode enriched scene context information compared to standard perspective image. Previous work has made lots of effort to solve the scene understanding task in a hybrid solution based on 2D-3D geometric reasoning, thus each sub-task is processed separately and few correlations are explored in this procedure. In this paper, we propose a fully 3D method for holistic indoor scene understanding which recovers the objects' shapes, oriented bounding boxes and the 3D room layout simultaneously from a single panorama. To maximize the exploration of the rich context information, we design a transformer-based context module to predict the representation and relationship among each component of the scene. In addition, we introduce a new dataset for scene understanding, including photo-realistic panoramas, high-fidelity depth images, accurately annotated room layouts, oriented object bounding boxes and shapes. Experiments on the synthetic and new datasets demonstrate that our method outperforms previous panoramic scene understanding methods in terms of both layout estimation and 3D object detection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- RAM: Recover Any 3D Human Motion in-the-WildSen Jia, Ning Zhu, Jinqin Zhong, Jiale Zhou 等CVPR 2026 · 被引用 12 次
- Diorama: Unleashing Zero-Shot Single-View 3D Indoor Scene ModelingQirui Wu, Denys Iliash, Daniel Ritchie, Manolis Savva 等ICCV 2025 · 被引用 4 次
- Pano3DComposer: Feed-Forward Compositional 3D Scene Generation from Single Panoramic ImageZidian Qiu, Ancong WuCVPR 2026 · 被引用 1 次
- PanSplat: 4K Panorama Synthesis with Feed-Forward Gaussian SplattingCheng Zhang, Haofei Xu, Qianyi Wu, Camilo Cruz Gambardella 等CVPR 2025
- Omnidirectional Multi-Object TrackingKai Luo, Hao Shi, Sheng Wu, Fei Teng 等CVPR 2025
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu 等ICCV 2021 · 被引用 2,397 次
相关 Paper
- DeepPanoContext: Panoramic 3D Scene Understanding with Holistic Scene Context Graph and Relation-based OptimizationCheng Zhang, Zhaopeng Cui, Cai Chen, Shuaicheng Liu 等ICCV 2021 · 被引用 43 次
- Total3DUnderstanding: Joint Layout, Object Pose and Mesh Reconstruction for Indoor Scenes From a Single ImageYinyu Nie, Xiaoguang Han, Shihui Guo, Yujian Zheng 等CVPR 2020
- Uni-3D: A Universal Model for Panoptic 3D Scene ReconstructionXiang Zhang, Zeyuan Chen, Fangyin Wei, Zhuowen TuICCV 2023 · 被引用 24 次
- Geometric Exploitation for Indoor Panoramic Semantic SegmentationDinh Duc Cao, Seok Joon Kim, Kyusung ChoNeurIPS 2024 · 被引用 14 次
- LED2-Net: Monocular 360deg Layout Estimation via Differentiable Depth RenderingFu-En Wang, Yu-Hsuan Yeh, Min Sun, Wei-Chen Chiu 等CVPR 2021
