DarSwin: Distortion Aware Radial Swin Transformer
Akshaya Athwale, Arman Afrasiyabi, Justin Lagüe, Ichrak Shili, Ola Ahmad, Jean-François Lalonde
Abstract
Wide-angle lenses are commonly used in perception tasks requiring a large field of view. Unfortunately, these lenses produce significant distortions making conventional models that ignore the distortion effects unable to adapt to wide-angle images. In this paper, we present a novel transformer-based model that automatically adapts to the distortion produced by wide-angle lenses. We leverage the physical characteristics of such lenses, which are analytically defined by the radial distortion profile (assumed to be known), to develop a distortion aware radial swin transformer (DarSwin). In contrast to conventional transformer-based architectures, DarSwin comprises a radial patch partitioning, a distortion-based sampling technique for creating token embeddings, and an angular position encoding for radial patch merging. We validate our method on classification tasks using synthetically distorted ImageNet data and show through extensive experiments that DarSwin can perform zero-shot adaptation to unseen distortions of different wide-angle lenses. Compared to other baselines, DarSwin achieves the best results (in terms of Top-1 accuracy) with significant gains when trained on bounded levels of distortions (very-low, low, medium, and high) and tested on all including out-of-distribution distortions. The code and models are publicly available at https://lvsn.github.io/darswin/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d29c8f0-1a45-4a67-8e38-a13a9103ae0bCited by top-tier papers1
Ask how each one uses itBuilds on7
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Vision Transformer with Deformable AttentionZhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li et al.CVPR 2022 · 835 citations
- WoodScape: A Multi-Task, Multi-Camera Fisheye Dataset for Autonomous DrivingSenthil Kumar Yogamani, Christian Witt, Hazem Rashed, Sanjaya Nayak et al.ICCV 2019 · 325 citations
- Bending Reality: Distortion-aware Transformers for Adapting to Panoramic Semantic SegmentationJiaming Zhang, Kailun Yang, Chaoxiang Ma, Simon Reiß et al.CVPR 2022 · 100 citations
Related papers
- RDCFace: Radial Distortion Correction for Face RecognitionHe Zhao, Xianghua Ying, Yongjie Shi, Xin Tong et al.CVPR 2020
- HEAL-SWIN: A Vision Transformer on the SphereOscar Carlsson, Jan E. Gerken, Hampus Linander, Heiner Spieß et al.CVPR 2024
- Semi-Supervised Wide-Angle Portraits Correction by Multi-Scale TransformerFushun Zhu, Shan Zhao, Peng Wang, Hao Wang et al.CVPR 2022 · 23 citations
- DaDA: Distortion-aware Domain Adaptation for Unsupervised Semantic SegmentationSujin Jang, Joohan Na, Dokwan OhNeurIPS 2022 · 13 citations
- Towards Complete Scene and Regular Shape for Distortion Rectification by Curve-Aware ExtrapolationKang Liao, Chunyu Lin, Yunchao Wei, Feng Li et al.ICCV 2021 · 9 citations
