AutoFocusFormer: Image Segmentation off the Grid
Ziwen Chen, Kaushik Patnaik, Shuangfei Zhai, Alvin Wan, Zhile Ren, Alexander G. Schwing, Alex Colburn, Fuxin Li
摘要
Real world images often have highly imbalanced content density. Some areas are very uniform, e.g., large patches of blue sky, while other areas are scattered with many small objects. Yet, the commonly used successive grid downsampling strategy in convolutional deep networks treats all areas equally. Hence, small objects are represented in very few spatial locations, leading to worse results in tasks such as segmentation. Intuitively, retaining more pixels representing small objects during downsampling helps to preserve important information. To achieve this, we propose AutoFocusFormer (AFF), a local-attention transformer image recognition backbone, which performs adaptive downsampling by learning to retain the most important pixels for the task. Since adaptive downsampling generates a set of pixels irregularly distributed on the image plane, we abandon the classic grid structure. Instead, we develop a novel point-based local attention block, facilitated by a balanced clustering module and a learnable neighborhood merging module, which yields representations for our point-based versions of state-of-the-art segmentation heads. Experiments show that our AutoFocusFormer (AFF) improves significantly over baseline models of similar sizes. * Work done while Chen Ziwen was an intern at Apple Inc. Image Remaining tokens Stage 2 Prediction AFF Swin Remaining tokens Stage 4 Comparison between on-grid model Swin and off-grid model AFF. AFF downsamples non automatically focusing on more textured, important image regions, and successfully captu the background. The red pixels indicate the locations of the remaining tokens.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
相关 Paper
- AdaFormer: Efficient Transformer with Adaptive Token Sparsification for Image Super-resolutionXiaotong Luo, Zekun Ai, Qiuyuan Liang, Ding Liu 等AAAI 2024 · 被引用 16 次
- SG-Former: Self-guided Transformer with Evolving Token ReallocationSucheng Ren, Xingyi Yang, Songhua Liu, Xinchao WangICCV 2023 · 被引用 70 次
- MixFormer: Mixing Features across Windows and DimensionsQiang Chen, Qiman Wu, Jian Wang, Qinghao Hu 等CVPR 2022 · 被引用 142 次
- Affine-Consistent Transformer for Multi-Class Cell Nuclei DetectionJunjia Huang, Haofeng Li, Xiang Wan, Guanbin LiICCV 2023 · 被引用 20 次
- Head-Free Lightweight Semantic Segmentation with Linear TransformerBo Dong, Pichao Wang, Fan WangAAAI 2023 · 被引用 129 次
