BiFormer: Vision Transformer with Bi-Level Routing Attention
Lei Zhu, Xinjiang Wang, Zhanghan Ke, Wayne Zhang, Rynson W. H. Lau
摘要
As the core building block of vision transformers, attention is a powerful tool to capture long-range dependency. However, such power comes at a cost: it incurs a huge computation burden and heavy memory footprint as pairwise token interaction across all spatial locations is computed. A series of works attempt to alleviate this problem by introducing handcrafted and content-agnostic sparsity into attention, such as restricting the attention operation to be inside local windows, axial stripes, or dilated windows. In contrast to these approaches, we propose a novel dynamic sparse attention via bi-level routing to enable a more flexible allocation of computations with content awareness. Specifically, for a query, irrelevant key-value pairs are first filtered out at a coarse region level, and then fine-grained token-to-token attention is applied in the union of remaining candidate regions (i.e., routed regions). We provide a simple yet effective implementation of the proposed bilevel routing attention, which utilizes the sparsity to save both computation and memory while involving only GPUfriendly dense matrix multiplications. Built with the proposed bi-level routing attention, a new general vision transformer, named BiFormer, is then presented. As BiFormer attends to a small subset of relevant tokens in a query adaptive manner without distraction from other irrelevant ones, it enjoys both good performance and high computational efficiency, especially in dense prediction tasks. Empirical results across several computer vision tasks such as image classification, object detection, and semantic segmentation verify the effectiveness of our design.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper56
- Rep ViT: Revisiting Mobile CNN From ViT PerspectiveAo Wang, Hui Chen, Zijia Lin, Jungong Han 等CVPR 2024 · 被引用 500 次
- Locality-Attending Vision TransformerSina Hajimiri, Farzad Beizaee, Fereshteh Shakeri, Christian Desrosiers 等ICLR 2026 · 被引用 427 次
- TransNeXt: Robust Foveal Visual Perception for Vision TransformersDai ShiCVPR 2024 · 被引用 313 次
- Demystify Mamba in Vision: A Linear Attention PerspectiveDongchen Han, Ziyi Wang, Zhuofan Xia, Yizeng Han 等NeurIPS 2024 · 被引用 287 次
- Multi-Scale VMamba: Hierarchy in Hierarchy Visual State Space ModelYuheng Shi, Minjing Dong, Chang XuNeurIPS 2024 · 被引用 129 次
它引用的顶会 Paper18
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
相关 Paper
- Enhancing Vision Transformers for Object Detection via Context-Aware Token Selection and PackingTianyi Zhang, Baoxin Li, Jae-sun Seo, Yu CaoICLR 2026
- Vision Transformer with Sparse Scan PriorYuguang Zhang, Qihang Fan, Huaibo HuangACM MM 2025 · 被引用 2 次
- Dynamic Grained Encoder for Vision TransformersLin Song, Songyang Zhang, Songtao Liu, Zeming Li 等NeurIPS 2021 · 被引用 41 次
- Orthogonal Transformer: An Efficient Vision Transformer Backbone with Token OrthogonalizationHuaibo Huang, Xiaoqiang Zhou, Ran HeNeurIPS 2022 · 被引用 35 次
- Sparsifiner: Learning Sparse Instance-Dependent Attention for Efficient Vision TransformersCong Wei, Brendan Duke, Ruowei Jiang, Parham Aarabi 等CVPR 2023
