One Meta-tuned Transformer is What You Need for Few-shot Learning
Xu Yang, Huaxiu Yao, Ying Wei
Abstract
Pre-trained vision transformers have revolutionized few-shot image classification, and it has been recently demonstrated that the previous common practice of meta-learning in synergy with these pre-trained transformers still holds significance. In this work, we design a new framework centered exclusively on attention mechanisms, called MetaFormer, which extends the vision transformers beyond patch token interactions to encompass relationships between samples and tasks simultaneously for further advancing their downstream task performance. Leveraging the intrinsical property of ViTs in handling local patch relationships, we propose Masked Sample Attention (MSA) to efficiently embed the sample relationships into the network, where an adaptive mask is attached for enhancing task-specific feature consistency and providing flexibility in switching between fewshot learning setups. To encapsulate task relationships while filtering out background noise, Patchgrained Task Attention (PTA) is designed to maintain a dynamic knowledge pool consolidating diverse patterns from historical tasks. MetaFormer demonstrates coherence and compatibility with off-the-shelf pre-trained vision transformers and shows significant improvements in both inductive and transductive few-shot learning scenarios, outperforming state-of-the-art methods by up to 8.77% and 6.25% on 12 in-domain and 10 crossdomain datasets, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on57
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- CAD: Co-Adapting Discriminative Features for Improved Few-Shot ClassificationPhilip Chikontwe, Soopil Kim, Sang Hyun ParkCVPR 2022 · 46 citations
- SpatialFormer: Semantic and Target Aware Attentions for Few-Shot LearningJinxiang Lai, Siqian Yang, Wenlong Wu, Tao Wu et al.AAAI 2023 · 21 citations
- Task-Adaptive Prompted Transformer for Cross-Domain Few-Shot LearningJiamin Wu, Xin Liu, Xiaotian Yin, Tianzhu Zhang et al.AAAI 2024 · 14 citations
- Focus Your Attention when Few-Shot ClassificationHaoqing Wang, Shibo Jie, Zhihong DengNeurIPS 2023 · 16 citations
- Attention Temperature Matters in ViT-Based Cross-Domain Few-Shot LearningYixiong Zou, Ran Ma, Yuhua Li, Ruixuan LiNeurIPS 2024 · 35 citations
