Prototype-Based Contrastive Learning with Stage-Wise Progressive Augmentation for Self-Supervised Fine-Grained Learning
Baofeng Tan, Xiu-Shen Wei, Lin Zhao
Abstract
In this paper, we mitigate the problem of Self-Supervised Learning (SSL) for fine-grained representation learning, aimed at distinguishing subtle differences within highly similar subordinate categories. Our preliminary analysis shows that SSL, especially the multi-stage alignment strategy, performs well on generic categories but struggles with finegrained distinctions. To overcome this limitation, we propose a prototype-based contrastive learning module with stagewise progressive augmentation. Unlike previous methods, our stage-wise progressive augmentation adapts data augmentation across stages to better suit SSL on fine-grained datasets. The prototype-based contrastive learning module captures both holistic and partial patterns, extracting global and local image representations to enhance feature discriminability. Experiments on popular fine-grained benchmarks for classification and retrieval tasks demonstrate the effectiveness of our method, and extensive ablation studies confirm the superiority of our proposals. Codes are available at https://github.com/SEU-VIPGroup/PAPN .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 399bb590-e68c-49af-b79c-20722df19315Cited by top-tier papers3
- Learning to See Through a Baby’s Eyes: Early Visual Diets Enable Robust Visual Intelligence in Humans and MachinesYusen Cai, Qing Lin, BHARGAVA SATYA NUNNA, Mengmi ZhangCVPR 2026 · 4 citations
- ArcDAE: Asymmetric Rectified Contrastive Diffusion Autoencoder for Unified Representation LearningGe Gao, Di Xiong, Zeke Xie, Jian Yang et al.ICML 2026
- Towards Fine-Grained Robustness: Attention-Guided Test-Time Prompt Tuning for Vision-Language ModelsJia-Wei Hai, Yijun Wang, Xiu-Shen WeiICML 2026
Builds on20
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
Related papers
- PaCL: Part-level Contrastive Learning for Fine-grained Few-shot Image ClassificationChuanming Wang, Huiyuan Fu, Huadong MaACM MM 2022 · 23 citations
- An Asymmetric Augmented Self-Supervised Learning Method for Unsupervised Fine-Grained Image HashingFeiran Hu, Chen-Lin Zhang, Jiangliang Guo, Xiu-Shen Wei et al.CVPR 2024
- Semantics-Consistent Feature Search for Self-Supervised Visual Representation LearningKaiyou Song, Shan Zhang, Zimeng Luo, Tong Wang et al.ICCV 2023 · 10 citations
- Part-level Semantic-guided Contrastive Learning for Fine-grained Visual ClassificationZhijian Lin, Hong HanICLR 2026
- Unleashing Potential of Unsupervised Pre-Training with Intra-Identity Regularization for Person Re-IdentificationZizheng Yang, Xin Jin, Kecheng Zheng, Feng ZhaoCVPR 2022 · 30 citations
