On the Role of Batch Size in Stochastic Conditional Gradient Methods
Rustem Islamov, Roman Machacek, Aurelien Lucchi, Antonio Silveti-Falls, Eduard Gorbunov, Volkan Cevher
摘要
We study the role of batch size in stochastic conditional gradient methods under a -Kurdyka–Łojasiewicz (-KL) condition. Focusing on momentum-based stochastic Frank–Wolfe–type conditional gradient algorithms (e.g., Scion), we derive a new analysis that explicitly captures the interaction between stepsize, batch size, and stochastic noise. Our study reveals a regime-dependent behavior: increasing the batch size initially improves optimization accuracy, but beyond a critical threshold, the benefits saturate and can eventually degrade performance under a fixed token budget. Notably, the theory predicts the magnitude of the optimal stepsize and aligns well with empirical practices observed in large-scale training. Leveraging these insights, we derive principled guidelines for selecting the batch size and stepsize, and propose an adaptive strategy that increases batch size and sequence length during training while preserving convergence guarantees. Preliminary experiments are consistent with the theoretical predictions and illustrate the emergence of the predicted scaling regimes. Overall, our results provide a theoretical framework for understanding batch-size scaling in stochastic conditional gradient methods and offer guidance for designing efficient training schedules in large-scale optimization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 被引用 598 次
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
- Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferGe Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor 等NeurIPS 2021 · 被引用 208 次
- Don't be lazy: CompleteP enables compute-efficient deep transformersNolan Dey, Bin Claire Zhang, Lorenzo Noci, Mufan Bill Li 等NeurIPS 2025 · 被引用 77 次
- Scalable Optimization in the Modular NormTim Large, Yang Liu, Jacob Huh, Hyojin Bahng 等NeurIPS 2024 · 被引用 70 次
相关 Paper
- Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed NoiseMaria-Eleni Sfyraki, Jun-Kun WangICML 2026 · 被引用 37 次
- Generalized Polyak Step Size for First Order Optimization with MomentumXiaoyu Wang, Mikael Johansson, Tong ZhangICML 2023 · 被引用 32 次
- Seesaw: Accelerating Training by Balancing Batch Size and Learning Rate SchedulingAlexandru Meterez, Depen Morwani, Jingfeng Wu, Costin-Andrei Oncescu 等ICLR 2026 · 被引用 3 次
- General Analysis of LMO-based Optimizers: Beyond Bounded VarianceEgor Shulgin, Mohamed Awad, Peter Richtarik, Eduard GorbunovICML 2026
- Critical Batch Size Revisited: A Simple Empirical Approach to Large-Batch Language Model TrainingWilliam Merrill, Shane Arora, Dirk Groeneveld, Hanna HajishirziNeurIPS 2025 · 被引用 23 次
