Neighborhood Attention Transformer
Ali Hassani, Steven Walton, Jiachen Li, Shen Li, Humphrey Shi
摘要
https://github.com/SHI-Labs/Neighborhood-Attention-Transformer Self Attention (ViT) Window Self Attention (Swin) Shifted Window Self Attention (Swin) Neighborhood Attention (NAT) Figure 1. An illustration of attention spans in Self Attention, (Shifted) Window Self Attention, and our Neighborhood Attention. Self Attention allows each token to attend to everything. Window Self Attention divides self attention into non-overlapping sub-windows, and is followed by Shifted Window Self Attention, which allows for out-of-window interactions that are necessary to receptive field expansion. Neighborhood Attention localizes attention to a neighborhood around each token, introducing local inductive biases, maintaining translational equivariance, and allowing receptive field growth without needing extra operations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper114
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
- Locality-Attending Vision TransformerSina Hajimiri, Farzad Beizaee, Fereshteh Shakeri, Christian Desrosiers 等ICLR 2026 · 被引用 427 次
- Demystify Mamba in Vision: A Linear Attention PerspectiveDongchen Han, Ziyi Wang, Zhuofan Xia, Yizeng Han 等NeurIPS 2024 · 被引用 287 次
- DiTFastAttn: Attention Compression for Diffusion Transformer ModelsZhihang Yuan, Hanling Zhang, Lu Pu, Xuefei Ning 等NeurIPS 2024 · 被引用 134 次
- Multi-Scale VMamba: Hierarchy in Hierarchy Visual State Space ModelYuheng Shi, Minjing Dong, Chang XuNeurIPS 2024 · 被引用 129 次
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
相关 Paper
- Slide-Transformer: Hierarchical Vision Transformer with Local Self-AttentionXuran Pan, Tianzhu Ye, Zhuofan Xia, Shiji Song 等CVPR 2023
- On the Connection between Local Attention and Dynamic Depth-wise ConvolutionQi Han, Zejia Fan, Qi Dai, Lei Sun 等ICLR 2022 · 被引用 144 次
- Learned Queries for Efficient Local AttentionMoab Arar, Ariel Shamir, Amit H. BermanoCVPR 2022 · 被引用 28 次
- Absolute Position Embedding Learns Sinusoid-like Waves for Attention Based on Relative PositionYuji Yamamoto, Takuya MatsuzakiEMNLP 2023 · 被引用 1 次
- Cross Aggregation Transformer for Image RestorationZheng Chen, Yulun Zhang, Jinjin Gu, Yongbing Zhang 等NeurIPS 2022 · 被引用 274 次
