Hydro: Surrogate-Based Hyperparameter Tuning Service in Datacenters
Qinghao Hu, Zhisheng Ye, Meng Zhang, Qiaoling Chen, Peng Sun, Yonggang Wen, Tianwei Zhang
Abstract
Hyperparameter tuning is an essential step in deep learning model development that provides better model performance at the cost of substantial resources. While existing systems can improve tuning efficiency, they still fail to handle large models with billions of parameters and efficiently leverage cluster resources. Motivated by these deficiencies, we introduce Hydro, a surrogate-based hyperparameter tuning service that optimizes tuning workloads in both the job-level and cluster-level granularities. Specifically, it consists of two key components: (1) Hydro Tuner automatically generates and optimizes surrogate models via scaling, parametrization and fusion; (2) Hydro Coordinator improves tuning efficiency and cluster-wide resource utilization by adaptively leveraging ephemeral and heterogeneous resources. Our comprehensive experiments on two tuning algorithms across six models show that Hydro Tuner can dramatically reduce tuning makespan by up to 78.5× compared with Ray Tune and no reduction in tuning quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d1ae97e2-3413-4335-921d-766990f2489eCited by top-tier papers8
- Characterization of Large Language Model Development in the DatacenterQinghao Hu, Zhisheng Ye, Zerui Wang, Guoteng Wang et al.NSDI 2024 · 192 citations
- Optimus: Accelerating Large-Scale Multi-Modal LLM Training by Bubble ExploitationWeiqi Feng, Yangrui Chen, Shaoyu Wang, Yanghua Peng et al.USENIX ATC 2025 · 34 citations
- Multiplexing Dynamic Deep Learning Workloads with SLO-awareness in GPU ClustersWenyan Chen, Chengzhi Lu, Huanle Xu, Kejiang Ye et al.EuroSys 2025 · 13 citations
- Reducing Energy Bloat in Large Model TrainingJae-Won Chung, Yile Gu, Insu Jang, Luoxi Meng et al.SOSP 2024 · 12 citations
- Hyperparameter Transfer Laws for Non-Recurrent Multi-Path Neural NetworksHaosong Zhang, Shenxi Wu, Xingjian Ma, Shirui Bian et al.ICML 2026 · 1 citation
Builds on30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley et al.SC 2021 · 576 citations
Related papers
- HUNTER: An Online Cloud Database Hybrid Tuning System for Personalized RequirementsBaoqing Cai, Yu Liu, Ce Zhang, Guangyu Zhang et al.SIGMOD 2022 · 52 citations
- RubberBand: cloud-based hyperparameter tuningUjval Misra, Richard Liaw, Lisa Dunlap, Romil Bhardwaj et al.EuroSys 2021 · 21 citations
- EasyScale: Elastic Training with Consistent Accuracy and Improved Utilization on GPUsMingzhen Li, Wencong Xiao, Hailong Yang, Biao Sun et al.SC 2023 · 16 citations
- Hyper-Tune: Towards Efficient Hyper-parameter Tuning at ScaleYang Li, Yu Shen, Huaijun Jiang, Wentao Zhang et al.VLDB 2022 · 32 citations
- A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via FaroBeomyeol Jeon, Chen Wang, Diana Arroyo, Alaa Youssef et al.EuroSys 2025
