EfficientTrain: Exploring Generalized Curriculum Learning for Training Visual Backbones
Yulin Wang, Yang Yue, Rui Lu, Tianjiao Liu, Zhao Zhong, Shiji Song, Gao Huang
Abstract
The superior performance of modern deep networks usually comes with a costly training procedure. This paper presents a new curriculum learning approach for the efficient training of visual backbones (e.g., vision Transformers). Our work is inspired by the inherent learning dynamics of deep networks: we experimentally show that at an earlier training stage, the model mainly learns to recognize some 'easier-to-learn' discriminative patterns within each example, e.g., the lower-frequency components of images and the original information before data augmentation. Driven by this phenomenon, we propose a curriculum where the model always leverages all the training data at each epoch, while the curriculum starts with only exposing the 'easier-to-learn' patterns of each example, and introduces gradually more difficult patterns. To implement this idea, we 1) introduce a cropping operation in the Fourier spectrum of the inputs, which enables the model to learn from only the lower-frequency components efficiently, 2) demonstrate that exposing the features of original images amounts to adopting weaker data augmentation, and 3) integrate 1) and 2) and design a curriculum learning schedule with a greedy-search algorithm. The resulting approach, EfficientTrain, is simple, general, yet surprisingly effective. As an off-the-shelf method, it reduces the wall-time training cost of a wide variety of popular models (e.g., ResNet, ConvNeXt, DeiT, PVT, Swin, and CSWin) by > 1.5× on ImageNet-1K/22K without sacrificing accuracy. It is also effective for self-supervised learning (e.g., MAE). Code is available at https://github. com/LeapLabTHU/EfficientTrain .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 204cd93b-991d-4adf-9b59-5ef245c492c1Cited by top-tier papers16
- FLatten Transformer: Vision Transformer using Focused Linear AttentionDongchen Han, Xuran Pan, Yizeng Han, Shiji Song et al.ICCV 2023 · 358 citations
- GSVA: Generalized Segmentation via Multimodal Large Language ModelsZhuofan Xia, Dongchen Han, Yizeng Han, Xuran Pan et al.CVPR 2024 · 42 citations
- Reusing Pretrained Models by Multi-linear Operators for Efficient TrainingYu Pan, Ye Yuan, Yichun Yin, Zenglin Xu et al.NeurIPS 2023 · 23 citations
- Advancing Open-Set Domain Generalization Using Evidential Bi-Level Hardest Domain SchedulerKunyu Peng, Di Wen, Kailun Yang, Ao Luo et al.NeurIPS 2024 · 20 citations
- Hardness-Aware Dynamic Curriculum Learning for Robust Multimodal Emotion Recognition with Missing ModalitiesRui Liu, Haolin Zuo, Zheng Lian, Hongyu Yuan et al.ACM MM 2025 · 6 citations
Builds on41
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
Related papers
- Automated Progressive Learning for Efficient Training of Vision TransformersChanglin Li, Bohan Zhuang, Guangrun Wang, Xiaodan Liang et al.CVPR 2022 · 28 citations
- A General and Efficient Training for Transformer via Token ExpansionWenxuan Huang, Yunhang Shen, Jiao Xie, Baochang Zhang et al.CVPR 2024
- UPDP: A Unified Progressive Depth Pruner for CNN and Vision TransformerJi Liu, Dehua Tang, Yuanxian Huang, Li Zhang et al.AAAI 2024 · 18 citations
- From Prototypes to General Distributions: An Efficient Curriculum for Masked Image ModelingJinhong Lin, Cheng-En Wu, Huanran Li, Jifan Zhang et al.CVPR 2025
- Efficient Pre-training of Masked Language Model via Concept-based Curriculum MaskingMingyu Lee, Jun-Hyung Park, Junho Kim, Kang-Min Kim et al.EMNLP 2022 · 8 citations
