DarSwin: Distortion Aware Radial Swin Transformer
Akshaya Athwale, Arman Afrasiyabi, Justin Lagüe, Ichrak Shili, Ola Ahmad, Jean-François Lalonde
摘要
Wide-angle lenses are commonly used in perception tasks requiring a large field of view. Unfortunately, these lenses produce significant distortions making conventional models that ignore the distortion effects unable to adapt to wide-angle images. In this paper, we present a novel transformer-based model that automatically adapts to the distortion produced by wide-angle lenses. We leverage the physical characteristics of such lenses, which are analytically defined by the radial distortion profile (assumed to be known), to develop a distortion aware radial swin transformer (DarSwin). In contrast to conventional transformer-based architectures, DarSwin comprises a radial patch partitioning, a distortion-based sampling technique for creating token embeddings, and an angular position encoding for radial patch merging. We validate our method on classification tasks using synthetically distorted ImageNet data and show through extensive experiments that DarSwin can perform zero-shot adaptation to unseen distortions of different wide-angle lenses. Compared to other baselines, DarSwin achieves the best results (in terms of Top-1 accuracy) with significant gains when trained on bounded levels of distortions (very-low, low, medium, and high) and tested on all including out-of-distribution distortions. The code and models are publicly available at https://lvsn.github.io/darswin/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Vision Transformer with Deformable AttentionZhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li 等CVPR 2022 · 被引用 835 次
- WoodScape: A Multi-Task, Multi-Camera Fisheye Dataset for Autonomous DrivingSenthil Kumar Yogamani, Christian Witt, Hazem Rashed, Sanjaya Nayak 等ICCV 2019 · 被引用 325 次
- Bending Reality: Distortion-aware Transformers for Adapting to Panoramic Semantic SegmentationJiaming Zhang, Kailun Yang, Chaoxiang Ma, Simon Reiß 等CVPR 2022 · 被引用 100 次
相关 Paper
- RDCFace: Radial Distortion Correction for Face RecognitionHe Zhao, Xianghua Ying, Yongjie Shi, Xin Tong 等CVPR 2020
- HEAL-SWIN: A Vision Transformer on the SphereOscar Carlsson, Jan E. Gerken, Hampus Linander, Heiner Spieß 等CVPR 2024
- Semi-Supervised Wide-Angle Portraits Correction by Multi-Scale TransformerFushun Zhu, Shan Zhao, Peng Wang, Hao Wang 等CVPR 2022 · 被引用 23 次
- DaDA: Distortion-aware Domain Adaptation for Unsupervised Semantic SegmentationSujin Jang, Joohan Na, Dokwan OhNeurIPS 2022 · 被引用 13 次
- Towards Complete Scene and Regular Shape for Distortion Rectification by Curve-Aware ExtrapolationKang Liao, Chunyu Lin, Yunchao Wei, Feng Li 等ICCV 2021 · 被引用 9 次
