Understanding the Training Speedup from Sampling with Approximate Losses
Rudrajit Das, Xi Chen, Bertram Ieong, Parikshit Bansal, Sujay Sanghavi
摘要
It is well known that selecting samples with large losses/gradients can significantly reduce the number of training steps. However, the selection overhead is often too high to yield any meaningful gains in terms of overall training time. In this work, we focus on the greedy approach of selecting samples with large approximate losses instead of exact losses in order to reduce the selection overhead. For smooth convex losses, we show that such a greedy strategy can converge to a constant factor of the minimum value of the average loss in fewer iterations than the standard approach of random selection. We also theoretically quantify the effect of the approximation level. We then develop SIFT which uses early exiting to obtain approximate losses with an intermediate layer's representations for sample selection. We evaluate SIFT on the task of training a 110M parameter 12 layer BERT base model, and show significant gains (in terms of training hours and number of backpropagation steps) without any optimized implementation over vanilla training. For e.g., to reach 64% validation accuracy, SIFT with exit at the first layer takes 43 hours compared to 57 hours of vanilla training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Are Greedy Task Orderings Better Than Random in Continual Linear Regression?Matan Tsipory, Ran Levinstein, Itay Evron, Mark Kong 等NeurIPS 2025 · 被引用 5 次
- Upweighting Easy Samples in Fine-Tuning Mitigates ForgettingSunny Sanyal, Hayden Prairie, Rudrajit Das, Ali Kavis 等ICML 2025
它引用的顶会 Paper4
- Confident Adaptive Language ModelingTal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani 等NeurIPS 2022 · 被引用 394 次
- A General Analysis of Example-Selection for Stochastic Gradient DescentYucheng Lu, Si Yi Meng, Christopher De SaICLR 2022 · 被引用 23 次
- The Right Tool for the Job: Matching Model and Instance ComplexitiesRoy Schwartz, Gabriel Stanovsky, Swabha Swayamdipta, Jesse Dodge 等ACL 2020 · 被引用 4 次
- Finding the SWEET Spot: Analysis and Improvement of Adaptive Inference in Low Resource SettingsDaniel Rotem, Michael Hassid, Jonathan Mamou, Roy SchwartzACL 2023 · 被引用 1 次
相关 Paper
- LeeBERT: Learned Early Exit for BERT with cross-level optimizationWei ZhuACL 2021
- Prioritized Training on Points that are Learnable, Worth Learning, and not yet LearntSören Mindermann, Jan Markus Brauner, Muhammed Razzak, Mrinank Sharma 等ICML 2022 · 被引用 237 次
- SkipBERT: Efficient Inference with Shallow Layer SkippingJue Wang, Ke Chen, Gang Chen, Lidan Shou 等ACL 2022
- AdapLeR: Speeding up Inference by Adaptive Length ReductionAli Modarressi, Hosein Mohebbi, Mohammad Taher PilehvarACL 2022 · 被引用 34 次
- EarlyBERT: Efficient BERT Training via Early-bird Lottery TicketsXiaohan Chen, Yu Cheng, Shuohang Wang, Zhe Gan 等ACL 2021
