Scattering Vision Transformer: Spectral Mixing Matters
Badri N. Patro, Vijay Agneeswaran
Abstract
Vision transformers have gained significant attention and achieved state-of-the-art performance in various computer vision tasks, including image classification, instance segmentation, and object detection. However, challenges remain in addressing attention complexity and effectively capturing fine-grained information within images. Existing solutions often resort to down-sampling operations, such as pooling, to reduce computational cost. Unfortunately, such operations are non-invertible and can result in information loss. In this paper, we present a novel approach called Scattering Vision Transformer (SVT) to tackle these challenges. SVT incorporates a spectrally scattering network that enables the capture of intricate image details. SVT overcomes the invertibility issue associated with down-sampling operations by separating low-frequency and high-frequency components. Furthermore, SVT introduces a unique spectral gating network utilizing Einstein multiplication for token and channel mixing, effectively reducing complexity. We show that SVT achieves state-of-the-art performance on the ImageNet dataset with a significant reduction in a number of parameters and FLOPS. SVT shows 2% improvement over LiTv2 and iFormer. SVT-H-S reaches 84.2% top-1 accuracy, while SVT-H-B reaches 85.2% (state-of-art for base versions) and SVT-H-L reaches 85.7% (again state-of-art for large versions). SVT also shows comparable results in other vision tasks such as instance segmentation. SVT also outperforms other transformers in transfer learning on standard datasets such as CIFAR10, CIFAR100, Oxford Flower, and Stanford Car datasets. The project page is available on this webpage.https://badripatro.github.io/svt/.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 49d3a534-e027-4a4d-8e77-5ce0b3e0e685Cited by top-tier papers7
- Graph Convolutions Enrich the Self-Attention in Transformers!Jeongwhan Choi, Hyowon Wi, Jayoung Kim, Yehjin Shin et al.NeurIPS 2024 · 24 citations
- Frequency-Dynamic Attention Modulation for Dense PredictionLinwei Chen, Lin Gu, Ying FuICCV 2025 · 13 citations
- Toward Diffusible High-Dimensional Latent Spaces: A Frequency PerspectiveBolin Lai, Xudong Wang, Saketh Rambhatla, James M. Rehg et al.CVPR 2026 · 7 citations
- Frequency-Aware Token Reduction for Efficient Vision TransformerDongJae Lee, Jiwan Hur, Jaehyun Choi, Jaemyung Yu et al.NeurIPS 2025 · 4 citations
- Multiple Feature Refining Network for Visual Emotion Distribution LearningQinfu Xu, Shaozu Yuan, Yiwei Wei, Jie Wu et al.AAAI 2025 · 2 citations
Related papers
- Vision Transformer with Sparse Scan PriorYuguang Zhang, Qihang Fan, Huaibo HuangACM MM 2025 · 2 citations
- Group Vision TransformerYaopeng Peng, Milan Sonka, Danny Z. ChenACM MM 2024
- Inception TransformerChenyang Si, Weihao Yu, Pan Zhou, Yichen Zhou et al.NeurIPS 2022 · 24 citations
- Focal Attention for Long-Range Interactions in Vision TransformersJianwei Yang, Chunyuan Li, Pengchuan Zhang, Xiyang Dai et al.NeurIPS 2021 · 228 citations
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu et al.ICCV 2021 · 2,397 citations
