Stochastic Optimization with Laggard Data Pipelines
Naman Agarwal, Rohan Anil, Tomer Koren, Kunal Talwar, Cyril Zhang
Abstract
State-of-the-art optimization is steadily shifting towards massively parallel pipelines with extremely large batch sizes. As a consequence, CPU-bound preprocessing and disk/memory/network operations have emerged as new performance bottlenecks, as opposed to hardware-accelerated gradient computations. In this regime, a recently proposed approach is data echoing (Choi et al., 2019) , which takes repeated gradient steps on the same batch while waiting for fresh data to arrive from upstream. We provide the first convergence analyses of "data-echoed" extensions of common optimization methods, showing that they exhibit provable improvements over their synchronous counterparts. Specifically, we show that in convex optimization with stochastic minibatches, data echoing affords speedups on the curvature-dominated part of the convergence rate, while maintaining the optimal statistical rate.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1491fae8-915c-4a73-9376-3a41f2f2fb71Cited by top-tier papers7
- Refurbish Your Training Data: Reusing Partially Augmented Samples for Faster Deep Neural Network TrainingGyewon Lee, Irene Lee, Hyeonmin Ha, Kyung-Geun Lee et al.USENIX ATC 2021 · 25 citations
- Algorithmic Instabilities of Accelerated Gradient DescentAmit Attia, Tomer KorenNeurIPS 2021 · 21 citations
- Acceleration via Fractal Learning Rate SchedulesNaman Agarwal, Surbhi Goel, Cyril ZhangICML 2021 · 19 citations
- Pecan: Cost-Efficient ML Data Preprocessing with Automatic Transformation Ordering and Hybrid PlacementDan Graur, Oto Mraz, Muyu Li, Mohammad Sepehr Pourghannad et al.USENIX ATC 2024 · 18 citations
- Progressive Compressed Records: Taking a Byte out of Deep Learning DataMichael Kuchnik, George Amvrosiadis, Virginia SmithVLDB 2021 · 11 citations
Builds on3
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 267 citations
- Extreme Tensoring for Low-Memory PreconditioningXinyi Chen, Naman Agarwal, Elad Hazan, Cyril Zhang et al.ICLR 2020 · 11 citations
Related papers
- Asynchronous SGD Beats Minibatch SGD Under Arbitrary DelaysKonstantin Mishchenko, Francis R. Bach, Mathieu Even, Blake E. WoodworthNeurIPS 2022 · 95 citations
- A General Analysis of Example-Selection for Stochastic Gradient DescentYucheng Lu, Si Yi Meng, Christopher De SaICLR 2022 · 23 citations
- Efficient Pipeline Planning for Expedited Distributed DNN TrainingZiyue Luo, Xiaodong Yi, Guoping Long, Shiqing Fan et al.INFOCOM 2022 · 19 citations
- Concurrent Adversarial Learning for Large-Batch TrainingYong Liu, Xiangning Chen, Minhao Cheng, Cho-Jui Hsieh et al.ICLR 2022 · 14 citations
- Minibatch vs Local SGD with Shuffling: Tight Convergence Bounds and BeyondChulhee Yun, Shashank Rajput, Suvrit SraICLR 2022 · 47 citations
