PanoSwin: a Pano-style Swin Transformer for Panorama Understanding
Zhixin Ling, Zhen Xing, Xiangdong Zhou, Manliang Cao, Guichun Zhou
Abstract
In panorama understanding, the widely used equirectangular projection (ERP) entails boundary discontinuity and spatial distortion. It severely deteriorates the conventional CNNs and vision Transformers on panoramas. In this paper, we propose a simple yet effective architecture named PanoSwin to learn panorama representations with ERP. To deal with the challenges brought by equirectangular projection, we explore a pano-style shift windowing scheme and novel pitch attention to address the boundary discontinuity and the spatial distortion, respectively. Besides, based on spherical distance and Cartesian coordinates, we adapt absolute positional embeddings and relative positional biases for panoramas to enhance panoramic geometry information. Realizing that planar image understanding might share some common knowledge with panorama understanding, we devise a novel two-stage learning framework to facilitate knowledge transfer from the planar images to panoramas. We conduct experiments against the state-ofthe-art on various panoramic tasks, i.e., panoramic object detection, panoramic classification, and panoramic layout estimation. The experimental results demonstrate the effectiveness of PanoSwin in panorama understanding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dfbd7f25-1cae-421f-9071-7309a594d0d3Cited by top-tier papers4
- TULIP: Transformer for Upsampling of LiDAR Point CloudsBin Yang, Patrick Pfreundschuh, Roland Siegwart, Marco Hutter et al.CVPR 2024 · 18 citations
- Fine-Grained Perception in Panoramic Scenes: A Novel Task, Dataset, and Method for Object Importance RankingJia Song, Chenglizhao Chen, Xu Yu, Shanchen PangAAAI 2025 · 1 citation
- Omnidirectional Multi-Object TrackingKai Luo, Hao Shi, Sheng Wu, Fei Teng et al.CVPR 2025
- SVFormer: Semi-supervised Video Transformer for Action RecognitionZhen Xing, Qi Dai, Han Hu, Jingjing Chen et al.CVPR 2023
Builds on12
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu et al.ICCV 2021 · 2,397 citations
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li et al.ICCV 2021 · 1,611 citations
Related papers
- PanelNet: Understanding 360 Indoor Environment via Panel RepresentationHaozheng Yu, Lu He, Bing Jian, Weiwei Feng et al.CVPR 2023
- Look at the Neighbor: Distortion-aware Unsupervised Domain Adaptation for Panoramic Semantic SegmentationXu Zheng, Tianbo Pan, Yunhao Luo, Lin WangICCV 2023 · 46 citations
- World-Shaper: A Unified Framework for 360° Panoramic EditingDong Liang, yuhao liu, Jinyuan Jia, Youjun Zhao et al.ICML 2026
- EGformer: Equirectangular Geometry-biased Transformer for 360 Depth EstimationIlwi Yun, Chanyong Shin, Hyunku Lee, Hyuk-Jae Lee et al.ICCV 2023 · 50 citations
- LGT-Net: Indoor Panoramic Room Layout Estimation with Geometry-Aware Transformer NetworkZhigang Jiang, Zhongzheng Xiang, Jinhua Xu, Ming ZhaoCVPR 2022 · 39 citations
