Three Guidelines You Should Know for Universally Slimmable Self-Supervised Learning
Yun-Hao Cao, Peiqin Sun, Shuchang Zhou
摘要
We propose universally slimmable self-supervised learning (dubbed as US3L) to achieve better accuracy-efficiency trade-offs for deploying self-supervised models across different devices. We observe that direct adaptation of selfsupervised learning (SSL) to universally slimmable networks misbehaves as the training process frequently collapses. We then discover that temporal consistent guidance is the key to the success of SSL for universally slimmable networks, and we propose three guidelines for the loss design to ensure this temporal consistency from a unified gradient perspective. Moreover, we propose dynamic sampling and group regularization strategies to simultaneously improve training efficiency and accuracy. Our US3L method has been empirically validated on both convolutional neural networks and vision transformers. With only once training and one copy of weights, our method outperforms various state-of-the-art methods (individually trained or not) on benchmarks including recognition, object detection and instance segmentation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Representing Part-Whole Hierarchies in Foundation Models by Learning Localizability, Composability, and Decomposability from Anatomy via Self-SupervisionMohammad Reza Hosseinzadeh Taher, Michael B. Gotway, Jianming LiangCVPR 2024 · 被引用 12 次
- Slicing Vision Transformer for Flexible InferenceYitian Zhang, Huseyin Coskun, Xu Ma, Huan Wang 等NeurIPS 2024 · 被引用 4 次
- Waxing-and-Waning: a Generic Similarity-based Framework for Efficient Self-Supervised LearningSheng Li, Chao Wu, Ao Li, Yanzhi Wang 等ICLR 2024 · 被引用 4 次
- Mutual Effort for Efficiency: A Similarity-based Token Pruning for Vision Transformers in Self-Supervised LearningSheng Li, Qitao Tan, Yue Dai, Zhenglun Kong 等ICLR 2025
它引用的顶会 Paper13
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 被引用 2,340 次
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 被引用 1,188 次
相关 Paper
- Task-Customized Self-Supervised Pre-training with Scalable Dynamic RoutingZhili Liu, Jianhua Han, Lanqing Hong, Hang Xu 等AAAI 2022 · 被引用 30 次
- UniVIP: A Unified Framework for Self-Supervised Visual Pre-trainingZhaowen Li, Yousong Zhu, Fan Yang, Wei Li 等CVPR 2022 · 被引用 29 次
- Effective Self-supervised Pre-training on Low-compute Networks without DistillationFuwen Tan, Fatemeh Sadat Saleh, Brais MartínezICLR 2023
- One-Shot Exemplars for Class Grounding in Self-Supervised LearningHaowen Cui, Shuo Chen, Jun Li, Jian YangICLR 2026
- Harnessing small projectors and multiple views for efficient vision pretrainingArna Ghosh, Kumar Krishna Agrawal, Shagun Sodhani, Adam Oberman 等NeurIPS 2024 · 被引用 5 次
