Curriculum Learning by Dynamic Instance Hardness
Tianyi Zhou, Shengjie Wang, Jeff A. Bilmes
Abstract
A good teacher can adjust a curriculum based on students' learning history. By analogy, in this paper, we study the dynamics of a deep neural network's (DNN) performance on individual samples during its learning process. The observed properties allow us to develop an adaptive curriculum that leads to faster learning of more accurate models. We introduce dynamic instance hardness (DIH), the exponential moving average of a sample's instantaneous hardness (e.g., a loss, or a change in output) over the training history. A low DIH indicates that a model retains knowledge about a sample over time. For DNNs, we find that a sample's DIH early in training predicts its DIH in later stages. Hence, we can train a model using samples mostly with higher DIH and safely deprioritize those with lower DIH. This motivates a DIH guided curriculum learning (DIHCL) procedure. Compared to existing CL methods: (1) DIH is more stable over time than using only instantaneous hardness, which is noisy due to stochastic training and DNN's non-smoothness; (2) DIHCL is computationally inexpensive since it uses only a byproduct of back-propagation and thus does not require extra inference. On 11 datasets, DIHCL significantly outperforms random mini-batch SGD and recent CL methods in terms of efficiency and final performance. The code of DIHCL is available at https://github.com/tianyizhou/DIHCL .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext abb66f04-5ff0-44f7-92ae-94d41fa85bbdCited by top-tier papers19
- Curriculum Disentangled Recommendation with Noisy Multi-feedbackHong Chen, Yudong Chen, Xin Wang, Ruobing Xie et al.NeurIPS 2021 · 88 citations
- EfficientTrain: Exploring Generalized Curriculum Learning for Training Visual BackbonesYulin Wang, Yang Yue, Rui Lu, Tianjiao Liu et al.ICCV 2023 · 39 citations
- TiDAL: Learning Training Dynamics for Active LearningSeong Min Kye, Kwanghee Choi, Hyeongmin Byun, Buru ChangICCV 2023 · 25 citations
- CUDA: Curriculum of Data Augmentation for Long-tailed RecognitionSumyeong Ahn, Jongwoo Ko, Se-Young YunICLR 2023 · 15 citations
- When Do Curricula Work in Federated Learning?Saeed Vahidian, Sreevatsank Kadaveru, Woonjoon Baek, Weijia Wang et al.ICCV 2023 · 12 citations
Builds on1
Related papers
- Adaptive Curriculum LearningYajing Kong, Liu Liu, Jun Wang, Dacheng TaoICCV 2021 · 61 citations
- Training Dynamics for Curriculum Learning: A Study on Monolingual and Cross-lingual NLUFenia Christopoulou, Gerasimos Lampouras, Ignacio IacobacciEMNLP 2022 · 3 citations
- SuperLoss: A Generic Loss for Robust Curriculum LearningThibault Castells, Philippe Weinzaepfel, Jérôme RevaudNeurIPS 2020 · 96 citations
- Batch Loss Score for Dynamic Data PruningQing Zhou, Bingxuan Zhao, Tao Yang, Hongyuan Zhang et al.CVPR 2026
- Curriculum Loss: Robust Learning and Generalization against Label CorruptionYueming Lyu, Ivor W. TsangICLR 2020 · 190 citations
