SPANet: Frequency-balancing Token Mixer using Spectral Pooling Aggregation Modulation
Guhnoo Yun, Juhan Yoo, Kijung Kim, Jeongho Lee, Dong Hwan Kim
摘要
Recent studies show that self-attentions behave like low-pass filters (as opposed to convolutions) and enhancing their high-pass filtering capability improves model performance. Contrary to this idea, we investigate existing convolution-based models with spectral analysis and observe that improving the low-pass filtering in convolution operations also leads to performance improvement. To account for this observation, we hypothesize that utilizing optimal token mixers that capture balanced representations of both high- and low-frequency components can enhance the performance of models. We verify this by decomposing visual features into the frequency domain and combining them in a balanced manner. To handle this, we replace the balancing problem with a mask filtering problem in the frequency domain. Then, we introduce a novel token-mixer named SPAM and leverage it to derive a MetaFormer model termed as SPANet. Experimental results show that the proposed method provides a way to achieve this balance, and the balanced representations of both high- and low-frequency components can improve the performance of models on multiple computer vision tasks. Our code is available at https://doranlyong.github.io/projects/spanet/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- ViT-CoMer: Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense PredictionsChunlong Xia, Xinliang Wang, Feng Lv, Xin Hao 等CVPR 2024 · 被引用 109 次
- Frequency-Dynamic Attention Modulation for Dense PredictionLinwei Chen, Lin Gu, Ying FuICCV 2025 · 被引用 13 次
- Svit-Split: Unleashing the Power of Vision Foundation Models Via Efficient Splitting HeadsYifan Li, Xin Li, Tianqin Li, Wenbin He 等ICCV 2025 · 被引用 2 次
- Frequency-Adaptive Dilated Convolution for Semantic SegmentationLinwei Chen, Lin Gu, Dezhi Zheng, Ying FuCVPR 2024
- Enabling True Global Perception in State Space Models for Visual TasksJie Hui, Zhenxiang Zhang, Wenyu Mi, Jianji WangICLR 2026
它引用的顶会 Paper38
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
相关 Paper
- FFT-Based Dynamic Token Mixer for VisionYuki Tatsunami, Masato TakiAAAI 2024 · 被引用 73 次
- Adaptive Frequency Filters As Efficient Global Token MixersZhipeng Huang, Zhizheng Zhang, Cuiling Lan, Zheng-Jun Zha 等ICCV 2023 · 被引用 96 次
- MetaFormer is Actually What You Need for VisionWeihao Yu, Mi Luo, Pan Zhou, Chenyang Si 等CVPR 2022 · 被引用 1,114 次
- Inception TransformerChenyang Si, Weihao Yu, Pan Zhou, Yichen Zhou 等NeurIPS 2022 · 被引用 24 次
- Learning A Sparse Transformer Network for Effective Image DerainingXiang Chen, Hao Li, Mingqiang Li, Jinshan PanCVPR 2023
