Three Guidelines You Should Know for Universally Slimmable Self-Supervised Learning
Yun-Hao Cao, Peiqin Sun, Shuchang Zhou
Abstract
We propose universally slimmable self-supervised learning (dubbed as US3L) to achieve better accuracy-efficiency trade-offs for deploying self-supervised models across different devices. We observe that direct adaptation of selfsupervised learning (SSL) to universally slimmable networks misbehaves as the training process frequently collapses. We then discover that temporal consistent guidance is the key to the success of SSL for universally slimmable networks, and we propose three guidelines for the loss design to ensure this temporal consistency from a unified gradient perspective. Moreover, we propose dynamic sampling and group regularization strategies to simultaneously improve training efficiency and accuracy. Our US3L method has been empirically validated on both convolutional neural networks and vision transformers. With only once training and one copy of weights, our method outperforms various state-of-the-art methods (individually trained or not) on benchmarks including recognition, object detection and instance segmentation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1c072b50-e0a9-41d3-a5d5-40a70b024f22Cited by top-tier papers4
- Representing Part-Whole Hierarchies in Foundation Models by Learning Localizability, Composability, and Decomposability from Anatomy via Self-SupervisionMohammad Reza Hosseinzadeh Taher, Michael B. Gotway, Jianming LiangCVPR 2024 · 12 citations
- Slicing Vision Transformer for Flexible InferenceYitian Zhang, Huseyin Coskun, Xu Ma, Huan Wang et al.NeurIPS 2024 · 4 citations
- Waxing-and-Waning: a Generic Similarity-based Framework for Efficient Self-Supervised LearningSheng Li, Chao Wu, Ao Li, Yanzhi Wang et al.ICLR 2024 · 4 citations
- Mutual Effort for Efficiency: A Similarity-based Token Pruning for Vision Transformers in Self-Supervised LearningSheng Li, Qitao Tan, Yue Dai, Zhenglun Kong et al.ICLR 2025
Builds on13
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 1,188 citations
Related papers
- Task-Customized Self-Supervised Pre-training with Scalable Dynamic RoutingZhili Liu, Jianhua Han, Lanqing Hong, Hang Xu et al.AAAI 2022 · 30 citations
- UniVIP: A Unified Framework for Self-Supervised Visual Pre-trainingZhaowen Li, Yousong Zhu, Fan Yang, Wei Li et al.CVPR 2022 · 29 citations
- Effective Self-supervised Pre-training on Low-compute Networks without DistillationFuwen Tan, Fatemeh Sadat Saleh, Brais MartínezICLR 2023
- One-Shot Exemplars for Class Grounding in Self-Supervised LearningHaowen Cui, Shuo Chen, Jun Li, Jian YangICLR 2026
- Harnessing small projectors and multiple views for efficient vision pretrainingArna Ghosh, Kumar Krishna Agrawal, Shagun Sodhani, Adam Oberman et al.NeurIPS 2024 · 5 citations
