ESPT: A Self-Supervised Episodic Spatial Pretext Task for Improving Few-Shot Learning
Yi Rong, Xiongbo Lu, Zhaoyang Sun, Yaxiong Chen, Shengwu Xiong
Abstract
Self-supervised learning (SSL) techniques have recently been integrated into the few-shot learning (FSL) framework and have shown promising results in improving the few-shot image classification performance. However, existing SSL approaches used in FSL typically seek the supervision signals from the global embedding of every single image. Therefore, during the episodic training of FSL, these methods cannot capture and fully utilize the local visual information in image samples and the data structure information of the whole episode, which are beneficial to FSL. To this end, we propose to augment the few-shot learning objective with a novel self-supervised Episodic Spatial Pretext Task (ESPT). Specifically, for each few-shot episode, we generate its corresponding transformed episode by applying a random geometric transformation to all the images in it. Based on these, our ESPT objective is defined as maximizing the local spatial relationship consistency between the original episode and the transformed one. With this definition, the ESPT-augmented FSL objective promotes learning more transferable feature representations that capture the local spatial features of different images and their inter-relational structural information in each input episode, thus enabling the model to generalize better to new categories with only a few samples. Extensive experiments indicate that our ESPT method achieves new state-of-the-art performance for few-shot image classification on three mainstay benchmark datasets. The source code will be available at: https://github.com/Whut-YiRong/ESPT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f200e86f-560b-46b8-8108-228cee8b2d3dCited by top-tier papers4
- VT-FSL: Bridging Vision and Text with LLMs for Few-Shot LearningWenhao Li, Qiangchang Wang, Xianjing Meng, Zhibin Wu et al.NeurIPS 2025 · 10 citations
- KNN Transformer with Pyramid Prompts for Few-Shot LearningWenhao Li, Qiangchang Wang, Peng Zhao, Yilong YinACM MM 2024 · 3 citations
- Revisiting Continuity of Image Tokens for Cross-domain Few-shot LearningShuai Yi, Yixiong Zou, Yuhua Li, Ruixuan LiICML 2025
- Manhattan Self-Attention Diffusion Residual Networks with Dynamic Bias Rectification for BCI-based Few-Shot LearningHao Wang, Li Xu, Yuntao Yu, Weiyue Ding et al.AAAI 2025
Builds on19
- Boosting Few-Shot Visual Learning With Self-SupervisionSpyros Gidaris, Andrei Bursuc, Nikos Komodakis, Patrick Pérez et al.ICCV 2019 · 445 citations
- CrossTransformers: spatially-aware few-shot transferCarl Doersch, Ankush Gupta, Andrew ZissermanNeurIPS 2020 · 420 citations
- Free Lunch for Few-shot Learning: Distribution CalibrationShuo Yang, Lu Liu, Min XuICLR 2021 · 378 citations
- Relational Embedding for Few-Shot ClassificationDahyun Kang, Heeseung Kwon, Juhong Min, Minsu ChoICCV 2021 · 254 citations
- Self-supervised Label Augmentation via Input TransformationsHankook Lee, Sung Ju Hwang, Jinwoo ShinICML 2020 · 218 citations
Related papers
- IEPT: Instance-Level and Episode-Level Pretext Tasks for Few-Shot LearningManli Zhang, Jianhong Zhang, Zhiwu Lu, Tao Xiang et al.ICLR 2021 · 103 citations
- Semantic Prompt for Few-Shot Image RecognitionCVPR 2023
- Exploring Complementary Strengths of Invariant and Equivariant Representations for Few-Shot LearningMamshad Nayeem Rizve, Salman H. Khan, Fahad Shahbaz Khan, Mubarak ShahCVPR 2021
- Pareto Self-Supervised Training for Few-Shot LearningZhengyu Chen, Jixie Ge, Heshen Zhan, Siteng Huang et al.CVPR 2021
- Few-Shot Learning With Global Class RepresentationsAoxue Li, Tiange Luo, Tao Xiang, Weiran Huang et al.ICCV 2019 · 119 citations
