GR-Gauge: Cost-efficient Training Configuration By Gauging the Gradient Redundancy
Guanjie Wang, Chen Chen
摘要
The recent success of artificial intelligence motivates many non-professional users to train their own models. Those users often resort to cloud training services, seeking to obtain a sufficiently accurate model at a modest cost, for which properly setting up the learning rate and batch size is crucial. While various Hyper-parameter Optimization (HPO) methods have been proposed in that regard, they largely act based on heavy-weight validation signals, being inefficient in the overall cost. We find that the model training process can be viewed as a two-dimensional voting process-with gradients for different iterations and from different samples; moreover, to attain cost-efficient training is to ensure that the gradient redundancy is within a proper range which is similar across diverse models. Based on that insight, we further introduce GR-Gauge, a general method that gauges the gradient redundancy to instruct HPO decisions like configuration searching and trial termination. Extensive experiments demonstrate that GR-Gauge can help attain near-optimal accuracy in much less time than existing methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep LearningAurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger 等OSDI 2021 · 被引用 258 次
- Zen-NAS: A Zero-Shot NAS for High-Performance Image RecognitionMing Lin, Pichao Wang, Zhenhong Sun, Hesen Chen 等ICCV 2021 · 被引用 164 次
- Layer-Wise Adaptive Model Aggregation for Scalable Federated LearningSunwoo Lee, Tuo Zhang, Amir Salman AvestimehrAAAI 2023 · 被引用 87 次
- Seizing Critical Learning Periods in Federated LearningGang Yan, Hao Wang, Jian LiAAAI 2022 · 被引用 51 次
相关 Paper
- BTTackler: A Diagnosis-based Framework for Efficient Deep Learning Hyperparameter OptimizationZhongyi Pei, Zhiyao Cen, Yipeng Huang, Chen Wang 等KDD 2024 · 被引用 1 次
- Frugal Optimization for Cost-related HyperparametersQingyun Wu, Chi Wang, Silu HuangAAAI 2021 · 被引用 51 次
- AUTOMATA: Gradient Based Data Subset Selection for Compute-Efficient Hyper-parameter TuningKrishnaTeja Killamsetty, Guttu Sai Abhishek, Aakriti, Ganesh Ramakrishnan 等NeurIPS 2022 · 被引用 37 次
- Hyperparameter Optimization Is Deceiving Us, and How to Stop ItA. Feder Cooper, Yucheng Lu, Jessica Zosa Forde, Christopher De SaNeurIPS 2021 · 被引用 40 次
- Scalable One-Pass Optimisation of High-Dimensional Weight-Update Hyperparameters by Implicit DifferentiationRoss M. Clarke, Elre Talea Oldewage, José Miguel Hernández-LobatoICLR 2022 · 被引用 9 次
