Resource-Guided Configuration Space Reduction for Deep Learning Models
Yanjie Gao, Yonghao Zhu, Hongyu Zhang, Haoxiang Lin, Mao Yang
Abstract
Deep learning models, like traditional software systems, provide a large number of configuration options. A deep learning model can be configured with different hyperparameters and neural architectures. Recently, AutoML (Automated Machine Learning) has been widely adopted to automate model training by systematically exploring diverse configurations. However, current AutoML approaches do not take into consideration the computational constraints imposed by various resources such as available memory, computing power of devices, or execution time. The training with non-conforming configurations could lead to many failed AutoML trial jobs or inappropriate models, which cause significant resource waste and severely slow down development productivity. In this paper, we propose DnnSAT, a resource-guided AutoML approach for deep learning models to help existing AutoML tools efficiently reduce the configuration space ahead of time. DnnSAT can speed up the search process and achieve equal or even better model learning performance because it excludes trial jobs not satisfying the constraints and saves resources for more trials. We formulate the resource-guided configuration space reduction as a constraint satisfaction problem. DnnSAT includes a unified analytic cost model to construct common constraints with respect to the model weight size, number of floating-point operations, model inference time, and GPU memory consumption. It then utilizes an SMT solver to obtain the satisfiable configurations of hyperparameters and neural architectures. Our evaluation results demonstrate the effectiveness of DnnSAT in accelerating stateof-the-art AutoML methods (Hyperparameter Optimization and Neural Architecture Search) with an average speedup from 1.19X to 3.95X on public benchmarks. We believe that DnnSAT can make AutoML more practical in a real-world environment with constrained resources. Index Terms-configurable systems, deep learning, AutoML, constraint solving * "FLOPS" denotes floating-point operations per second, and "SLA" stands for service-level agreement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4dcc8aac-5dfa-4711-9f70-4c68e35cab2eCited by top-tier papers6
- An Empirical Study on Low GPU Utilization of Deep Learning JobsYanjie Gao, Yichen He, Xinze Li, Bo Zhao et al.ICSE 2024 · 22 citations
- Adapting Multi-objectivized Software Configuration TuningTao Chen, Miqing LiFSE 2024 · 14 citations
- FedSlice: Protecting Federated Learning Models from Malicious Participants with Model SlicingZiqi Zhang, Yuanchun Li, Bingyan Liu, Yifeng Cai et al.ICSE 2023 · 8 citations
- Distilled Lifelong Self-Adaptation for Configurable SystemsYulong Ye, Tao Chen, Miqing LiICSE 2025 · 7 citations
- REFTY: Refinement Types for Valid Deep Learning ModelsYanjie Gao, Zhengxian Li, Haoxiang Lin, Hongyu Zhang et al.ICSE 2022 · 4 citations
Builds on3
- An empirical study on program failures of deep learning jobsRu Zhang, Wencong Xiao, Hongyu Zhang, Yu Liu et al.ICSE 2020 · 96 citations
- Budgeted Training: Rethinking Deep Neural Network Training Under Resource ConstraintsMengtian Li, Ersin Yumer, Deva RamananICLR 2020 · 58 citations
- AMS: generating AutoML search spaces from weak specificationsJosé Pablo Cambronero, Jürgen Cito, Martin C. RinardFSE 2020 · 12 citations
Related papers
- Saturn: An Optimized Data System for Multi-Large-Model Deep Learning WorkloadsKabir Nagrecha, Arun KumarVLDB 2024 · 8 citations
- CoSA: Scheduling by Constrained Optimization for Spatial AcceleratorsQijing Huang, Aravind Kalaiah, Minwoo Kang, James Demmel et al.ISCA 2021 · 120 citations
- On EDA-Driven Learning for SAT SolvingMin Li, Zhengyuan Shi, Qiuxia Lai, Sadaf Khan et al.DAC 2023 · 4 citations
- SmartMoE: Efficiently Training Sparsely-Activated Models through Combining Offline and Online ParallelizationMingshu Zhai, Jiaao He, Zixuan Ma, Zan Zong et al.USENIX ATC 2023 · 96 citations
- NeuroBack: Improving CDCL SAT Solving using Graph Neural NetworksWenxi Wang, Yang Hu, Mohit Tiwari, Sarfraz Khurshid et al.ICLR 2024 · 25 citations
