Addressing Exacerbated Attention Sink for Source-Free Cross-Domain Few-Shot Learning
Shuai Yi, Yixiong Zou, Yuhua Li, Ruixuan Li
Abstract
Vision-language models (VLMs) like CLIP have shown impressive generalization capabilities, yet their potential for Cross-Domain Few-Shot Learning (CDFSL) remains underexplored, where the model needs to transfer source-domain information to target domains with scarce training data. While the attention sink phenomenon has been observed in VLMs for certain tasks, its role in CDFSL scenarios has not been studied. In this paper, we uncover a critical issue overlooked by prior works: standard target-domain few-shot fine-tuning in CDFSL significantly exacerbates the attention sink problem, leading to poor discriminability across classes. To understand this phenomenon, through extensive experiments, we interpret it as the model's shortcut learning for domain adaptation: to overcome the huge domain gap between the source and target domains, the model shows a high tendency to push tokens that are initially closer to target-domain classes (i.e., simple tokens) to be even closer to these classes, exacerbating the attention sink and wasting the capability of learning other discriminative but initially further tokens (i.e., hard tokens). To address this, we propose a novel approach to dynamically re-weight tokens according to their relevance with target-domain classes during the target-domain finetuning, which explicitly suppresses the model's reliance on these simple tokens and enhances the learning of hard tokens, reducing sink tokens and enhancing discriminability. Extensive experiments on four benchmark datasets validate the rationale of our method, demonstrating new state-of-the-art performance. Our codes are available at https://github.com/shuaiyi308/TIR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b4f3d173-fb83-4dc1-967e-ddbe522fbca5Cited by top-tier papers1
Ask how each one uses itBuilds on27
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 769 citations
- Perception Encoder: The best visual embeddings are not at the output of the networkDaniel Bolya, Po-Yao Huang, Peize Sun, Jang Hyun Cho et al.NeurIPS 2025 · 359 citations
Related papers
- Mind the Discriminability Trap in Source-Free Cross-domain Few-shot LearningZhenyu Zhang, Yixiong Zou, Yuhua Li, Ruixuan Li et al.CVPR 2026 · 6 citations
- Interpretable Cross-Domain Few-Shot Learning with Rectified Target-Domain Local AlignmentYaze Zhao, Yixiong Zou, Yuhua Li, Ruixuan LiCVPR 2026 · 5 citations
- Attention Temperature Matters in ViT-Based Cross-Domain Few-Shot LearningYixiong Zou, Ran Ma, Yuhua Li, Ruixuan LiNeurIPS 2024 · 35 citations
- Adaptive Parameter Selection for Tuning Vision-Language ModelsYi Zhang, Yi-Xuan Deng, Meng-Hao Guo, Shi-Min HuCVPR 2025
- Random Registers for Cross-Domain Few-Shot LearningShuai Yi, Yixiong Zou, Yuhua Li, Ruixuan LiICML 2025
