Understanding the Training Speedup from Sampling with Approximate Losses
Rudrajit Das, Xi Chen, Bertram Ieong, Parikshit Bansal, Sujay Sanghavi
Abstract
It is well known that selecting samples with large losses/gradients can significantly reduce the number of training steps. However, the selection overhead is often too high to yield any meaningful gains in terms of overall training time. In this work, we focus on the greedy approach of selecting samples with large approximate losses instead of exact losses in order to reduce the selection overhead. For smooth convex losses, we show that such a greedy strategy can converge to a constant factor of the minimum value of the average loss in fewer iterations than the standard approach of random selection. We also theoretically quantify the effect of the approximation level. We then develop SIFT which uses early exiting to obtain approximate losses with an intermediate layer's representations for sample selection. We evaluate SIFT on the task of training a 110M parameter 12 layer BERT base model, and show significant gains (in terms of training hours and number of backpropagation steps) without any optimized implementation over vanilla training. For e.g., to reach 64% validation accuracy, SIFT with exit at the first layer takes 43 hours compared to 57 hours of vanilla training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Are Greedy Task Orderings Better Than Random in Continual Linear Regression?Matan Tsipory, Ran Levinstein, Itay Evron, Mark Kong et al.NeurIPS 2025 · 5 citations
- Upweighting Easy Samples in Fine-Tuning Mitigates ForgettingSunny Sanyal, Hayden Prairie, Rudrajit Das, Ali Kavis et al.ICML 2025
Builds on4
- Confident Adaptive Language ModelingTal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani et al.NeurIPS 2022 · 394 citations
- A General Analysis of Example-Selection for Stochastic Gradient DescentYucheng Lu, Si Yi Meng, Christopher De SaICLR 2022 · 23 citations
- The Right Tool for the Job: Matching Model and Instance ComplexitiesRoy Schwartz, Gabriel Stanovsky, Swabha Swayamdipta, Jesse Dodge et al.ACL 2020 · 4 citations
- Finding the SWEET Spot: Analysis and Improvement of Adaptive Inference in Low Resource SettingsDaniel Rotem, Michael Hassid, Jonathan Mamou, Roy SchwartzACL 2023 · 1 citation
Related papers
- LeeBERT: Learned Early Exit for BERT with cross-level optimizationWei ZhuACL 2021
- Prioritized Training on Points that are Learnable, Worth Learning, and not yet LearntSören Mindermann, Jan Markus Brauner, Muhammed Razzak, Mrinank Sharma et al.ICML 2022 · 237 citations
- SkipBERT: Efficient Inference with Shallow Layer SkippingJue Wang, Ke Chen, Gang Chen, Lidan Shou et al.ACL 2022
- AdapLeR: Speeding up Inference by Adaptive Length ReductionAli Modarressi, Hosein Mohebbi, Mohammad Taher PilehvarACL 2022 · 34 citations
- EarlyBERT: Efficient BERT Training via Early-bird Lottery TicketsXiaohan Chen, Yu Cheng, Shuohang Wang, Zhe Gan et al.ACL 2021
