REViT: Roto-reflection Equivariant Convolutional Vision Transformer
Sheir A. Zaheer, Alexander Holston, Chan Youn Park
Abstract
In this paper, we propose a discrete roto-reflection group equivariant vision transformer with convolutional attention. Roto-reflection equivariant networks preserve the rotational, flip and positional symmetry in feature maps, making them useful for tasks where orientation of the inputs is relevant to the model outputs. In image classification and object detection, most of the studies on roto-reflection equivariant models have focused on using convolutional neural networks rather than vision transformers. In this paper, we examine the challenges involved in achieving equivariance in vision transformers, and we propose a simpler way to implement a discretized roto-reflection group equivariant vision transformer. The experimental results demonstrate that our approach outperforms the existing approaches for developing discrete roto-reflection group equivariant neural networks for image classification.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ff31f99-0e48-4826-8efb-3d1b69922023Builds on19
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu et al.ICCV 2021 · 2,397 citations
- E(n) Equivariant Graph Neural NetworksVictor Garcia Satorras, Emiel Hoogeboom, Max WellingICML 2021 · 1,432 citations
- Equivariant message passing for the prediction of tensorial properties and molecular spectraKristof Schütt, Oliver T. Unke, Michael GasteggerICML 2021 · 736 citations
- Rethinking and Improving Relative Position Encoding for Vision TransformerKan Wu, Houwen Peng, Minghao Chen, Jianlong Fu et al.ICCV 2021 · 427 citations
Related papers
- Co-Attentive Equivariant Neural Networks: Focusing Equivariance On Transformations Co-Occurring in DataDavid W. Romero, Mark HoogendoornICLR 2020 · 24 citations
- Group Equivariant Stand-Alone Self-Attention For VisionDavid W. Romero, Jean-Baptiste CordonnierICLR 2021 · 72 citations
- Reflection and Rotation Symmetry Detection via Equivariant LearningAhyun Seo, Byungjin Kim, Suha Kwak, Minsu ChoCVPR 2022 · 12 citations
- Learning Partial Equivariances From DataDavid W. Romero, Suhas LohitNeurIPS 2022 · 54 citations
- Attentive Group Equivariant Convolutional NetworksDavid W. Romero, Erik J. Bekkers, Jakub M. Tomczak, Mark HoogendoornICML 2020 · 99 citations
