Revisiting Continuity of Image Tokens for Cross-domain Few-shot Learning
Shuai Yi, Yixiong Zou, Yuhua Li, Ruixuan Li
Abstract
Vision Transformer (ViT) has achieved remarkable success due to its large-scale pretraining on general domains, but it still faces challenges when applying it to downstream distant domains that have only scarce training data, which gives rise to the Cross-Domain Few-Shot Learning (CDFSL) task. Inspired by Self-Attention's insensitivity to token orders, we find an interesting phenomenon neglected in current works: disrupting the continuity of image tokens (i.e., making pixels not smoothly transited across patches) in ViT leads to a noticeable performance decline in the general (source) domain but only a marginal decrease in downstream target domains. This questions the role of image tokens' continuity in ViT's generalization under large domain gaps. In this paper, we delve into this phenomenon for an interpretation. We find continuity aids ViT in learning larger spatial patterns, which are harder to transfer than smaller ones, enlarging domain distances. Meanwhile, it implies that only smaller patterns within each patch could be transferred under extreme domain gaps. Based on this interpretation, we further propose a simple yet effective method for CDFSL that better disrupts the continuity of image tokens, encouraging the model to rely less on large patterns and more on smaller ones. Extensive experiments show the effectiveness of our method in reducing domain gaps and outperforming state-of-the-art works. Codes and models are available at https://github.com/shuaiyi308/ReCIT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 057e6ff1-aecf-45b9-bcdd-7336300ae4b1Cited by top-tier papers3
- Addressing Exacerbated Attention Sink for Source-Free Cross-Domain Few-Shot LearningShuai Yi, Yixiong Zou, Yuhua Li, Ruixuan LiCVPR 2026 · 3 citations
- Improving CLIP Adaptation by Breaking Tail Alignment for Source-Free Cross-Domain Few-Shot LearningShuai Yi, Yixiong Zou, Yuhua Li, Ruixuan LiICML 2026 · 1 citation
- Language Does Matter for Cross-Domain Few-Shot Visual Feature EnhancementFei Zhou, Xiwen Zhang, Qingqing Qiu, Lei Zhang et al.CVPR 2026
Builds on25
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
- Intriguing Properties of Vision TransformersMuzammal Naseer, Kanchana Ranasinghe, Salman Khan, Munawar Hayat et al.NeurIPS 2021 · 863 citations
Related papers
- Attention Temperature Matters in ViT-Based Cross-Domain Few-Shot LearningYixiong Zou, Ran Ma, Yuhua Li, Ruixuan LiNeurIPS 2024 · 35 citations
- A Closer Look at the CLS Token for Cross-Domain Few-Shot LearningYixiong Zou, Shuai Yi, Yuhua Li, Ruixuan LiNeurIPS 2024 · 40 citations
- Random Registers for Cross-Domain Few-Shot LearningShuai Yi, Yixiong Zou, Yuhua Li, Ruixuan LiICML 2025
- EViT: Expediting Vision Transformers via Token ReorganizationsYouwei Liang, Chongjian Ge, Zhan Tong, Yibing Song et al.ICLR 2022 · 137 citations
- Reconstruction Target Matters in Masked Image Modeling for Cross-Domain Few-Shot LearningRan Ma, Yixiong Zou, Yuhua Li, Ruixuan LiAAAI 2025 · 2 citations
