Bag of Instances Aggregation Boosts Self-supervised Distillation
Haohang Xu, Jiemin Fang, Xiaopeng Zhang, Lingxi Xie, Xinggang Wang, Wenrui Dai, Hongkai Xiong, Qi Tian
Abstract
Recent advances in self-supervised learning have experienced remarkable progress, especially for contrastive learning based methods, which regard each image as well as its augmentations as an individual class and try to distinguish them from all other images. However, due to the large quantity of exemplars, this kind of pretext task intrinsically suffers from slow convergence and is hard for optimization. This is especially true for small scale models, which we find the performance drops dramatically comparing with its supervised counterpart. In this paper, we propose a simple but effective distillation strategy for unsupervised learning. The highlight is that the relationship among similar samples counts and can be seamlessly transferred to the student to boost the performance. Our method, termed as BINGO, which is short for Bag of InstaNces aGgregatiOn, targets at transferring the relationship learned by the teacher to the student. Here bag of instances indicates a set of similar samples constructed by the teacher and are grouped within a bag, and the goal of distillation is to aggregate compact representations over the student with respect to instances in a bag. Notably, BINGO achieves new state-of-the-art performance on small scale models, i.e., 65.5% and 68.9% top-1 accuracies with linear evaluation on ImageNet, using ResNet-18 and ResNet-34 as backbone, respectively, surpassing baselines (52.5% and 57.4% top-1 accuracies) by a significant margin. The code is available at https://github.com/haohang96/bingo . * Equal contributions. The work was done during the internship of Haohang Xu and Jiemin Fang at Huawei Inc.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Dual Contrastive Learning for Spatio-temporal RepresentationShuangrui Ding, Rui Qian, Hongkai XiongACM MM 2022 · 20 citations
- Pixel-Wise Contrastive DistillationJunqiang Huang, Zichao GuoICCV 2023 · 8 citations
- Learning Task-Agnostic Representations through Multi-Teacher DistillationPhilippe Formont, Maxime Darrin, Banafsheh Karimian, Eric Granger et al.NeurIPS 2025 · 6 citations
- New Insights for the Stability-Plasticity Dilemma in Online Continual LearningDahuin Jung, Dongjin Lee, Sunwon Hong, Hyemi Jang et al.ICLR 2023 · 4 citations
- Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representation LearningRui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen et al.CVPR 2023
Builds on12
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
Related papers
- SEED: Self-supervised Distillation For Visual RepresentationZhiyuan Fang, Jianfeng Wang, Lijuan Wang, Lei Zhang et al.ICLR 2021 · 213 citations
- On the Efficacy of Small Self-Supervised Contrastive Models without Distillation SignalsHaizhou Shi, Youcai Zhang, Siliang Tang, Wenjie Zhu et al.AAAI 2022 · 17 citations
- Unsupervised Representation Transfer for Small Networks: I Believe I Can Distill On-the-FlyHee Min Choi, Hyoa Kang, Dokwan OhNeurIPS 2021 · 13 citations
- CompRess: Self-Supervised Learning by Compressing RepresentationsSoroush Abbasi Koohpayegani, Ajinkya Tejankar, Hamed PirsiavashNeurIPS 2020 · 105 citations
- Slimmed Asymmetrical Contrastive Learning and Cross Distillation for Lightweight Model TrainingJian Meng, Li Yang, Kyungmin Lee, Jinwoo Shin et al.NeurIPS 2023 · 3 citations
