PanoSwin: a Pano-style Swin Transformer for Panorama Understanding
Zhixin Ling, Zhen Xing, Xiangdong Zhou, Manliang Cao, Guichun Zhou
摘要
In panorama understanding, the widely used equirectangular projection (ERP) entails boundary discontinuity and spatial distortion. It severely deteriorates the conventional CNNs and vision Transformers on panoramas. In this paper, we propose a simple yet effective architecture named PanoSwin to learn panorama representations with ERP. To deal with the challenges brought by equirectangular projection, we explore a pano-style shift windowing scheme and novel pitch attention to address the boundary discontinuity and the spatial distortion, respectively. Besides, based on spherical distance and Cartesian coordinates, we adapt absolute positional embeddings and relative positional biases for panoramas to enhance panoramic geometry information. Realizing that planar image understanding might share some common knowledge with panorama understanding, we devise a novel two-stage learning framework to facilitate knowledge transfer from the planar images to panoramas. We conduct experiments against the state-ofthe-art on various panoramic tasks, i.e., panoramic object detection, panoramic classification, and panoramic layout estimation. The experimental results demonstrate the effectiveness of PanoSwin in panorama understanding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- TULIP: Transformer for Upsampling of LiDAR Point CloudsBin Yang, Patrick Pfreundschuh, Roland Siegwart, Marco Hutter 等CVPR 2024 · 被引用 18 次
- Fine-Grained Perception in Panoramic Scenes: A Novel Task, Dataset, and Method for Object Importance RankingJia Song, Chenglizhao Chen, Xu Yu, Shanchen PangAAAI 2025 · 被引用 1 次
- Omnidirectional Multi-Object TrackingKai Luo, Hao Shi, Sheng Wu, Fei Teng 等CVPR 2025
- SVFormer: Semi-supervised Video Transformer for Action RecognitionZhen Xing, Qi Dai, Han Hu, Jingjing Chen 等CVPR 2023
它引用的顶会 Paper12
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu 等ICCV 2021 · 被引用 2,397 次
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li 等ICCV 2021 · 被引用 1,611 次
相关 Paper
- PanelNet: Understanding 360 Indoor Environment via Panel RepresentationHaozheng Yu, Lu He, Bing Jian, Weiwei Feng 等CVPR 2023
- Look at the Neighbor: Distortion-aware Unsupervised Domain Adaptation for Panoramic Semantic SegmentationXu Zheng, Tianbo Pan, Yunhao Luo, Lin WangICCV 2023 · 被引用 46 次
- World-Shaper: A Unified Framework for 360° Panoramic EditingDong Liang, yuhao liu, Jinyuan Jia, Youjun Zhao 等ICML 2026
- EGformer: Equirectangular Geometry-biased Transformer for 360 Depth EstimationIlwi Yun, Chanyong Shin, Hyunku Lee, Hyuk-Jae Lee 等ICCV 2023 · 被引用 50 次
- LGT-Net: Indoor Panoramic Room Layout Estimation with Geometry-Aware Transformer NetworkZhigang Jiang, Zhongzheng Xiang, Jinhua Xu, Ming ZhaoCVPR 2022 · 被引用 39 次
