BTTackler: A Diagnosis-based Framework for Efficient Deep Learning Hyperparameter Optimization
Zhongyi Pei, Zhiyao Cen, Yipeng Huang, Chen Wang, Lin Liu, Philip S. Yu, Mingsheng Long, Jianmin Wang
摘要
Hyperparameter optimization (HPO) is known to be costly in deep learning, especially when leveraging automated approaches. Most of the existing automated HPO methods are accuracy-based, i.e., accuracy metrics are used to guide the trials of different hyperparameter configurations amongst a specific search space. However, many trials may encounter severe training problems, such as vanishing gradients and insufficient convergence, which can hardly be reflected by accuracy metrics in the early stages of the training and often result in poor performance. This leads to an inefficient optimization trajectory because the bad trials occupy considerable computation resources and reduce the probability of finding excellent hyperparameter configurations within a time limitation. In this paper, we propose Bad Trial Tackler (BTTackler), a novel HPO framework that introduces training diagnosis to identify training problems automatically and hence tackles bad trials. BTTackler diagnoses each trial by calculating a set of carefully designed quantified indicators and triggers early termination if any training problems are detected. Evaluations are performed on representative HPO tasks consisting of three classical deep neural networks (DNN) and four widely used HPO methods. To better quantify the effectiveness of an automated HPO method, we propose two new measurements based on accuracy and time consumption. Results show the advantage of BTTackler on two-fold: (1) it reduces 40.33% of time consumption to achieve the same accuracy comparable to baseline methods on average and (2) it conducts 44.5% more top-10 trials than baseline methods on average within a given time budget. We also released an open-source Python library that allows users to easily apply BTTackler to automated HPO processes with minimal code changes://github.com/thuml/BTTackler.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- MTSClean: Efficient Constraint-based Cleaning for Multi-Dimensional Time Series DataXiaoou Ding, Yichen Song, Hongzhi Wang, Chen Wang 等VLDB 2024 · 被引用 9 次
- GR-Gauge: Cost-efficient Training Configuration By Gauging the Gradient RedundancyGuanjie Wang, Chen ChenCVPR 2026
它引用的顶会 Paper8
- DeepLocalize: Fault Localization for Deep Neural NetworksMohammad Wardat, Wei Le, Hridesh RajanICSE 2021 · 被引用 93 次
- Detecting numerical bugs in neural network architecturesYuhao Zhang, Luyao Ren, Liqian Chen, Yingfei Xiong 等FSE 2020 · 被引用 66 次
- AUTOTRAINER: An Automatic DNN Training Problem Detection and Repair SystemXiaoyu Zhang, Juan Zhai, Shiqing Ma, Chao ShenICSE 2021 · 被引用 62 次
- UMLAUT: Debugging Deep Learning Programs using Program Structure and Model BehaviorEldon Schoop, Forrest Huang, Bjoern HartmannCHI 2021 · 被引用 50 次
- DeepDiagnosis: Automatically Diagnosing Faults and Recommending Actionable Fixes in Deep Learning ProgramsMohammad Wardat, Breno Dantas Cruz, Wei Le, Hridesh RajanICSE 2022 · 被引用 46 次
相关 Paper
- Hippo: Sharing Computations in Hyper-Parameter OptimizationAhnjae Shin, Joo Seong Jeong, Do Yoon Kim, Soyoung Jung 等VLDB 2022 · 被引用 6 次
- AUTOMATA: Gradient Based Data Subset Selection for Compute-Efficient Hyper-parameter TuningKrishnaTeja Killamsetty, Guttu Sai Abhishek, Aakriti, Ganesh Ramakrishnan 等NeurIPS 2022 · 被引用 37 次
- HyperSTAR: Task-Aware Hyperparameters for Deep NetworksGaurav Mittal, Chang Liu, Nikolaos Karianakis, Victor Fragoso 等CVPR 2020
- A New Linear Scaling Rule for Private Adaptive Hyperparameter OptimizationAshwinee Panda, Xinyu Tang, Saeed Mahloujifar, Vikash Sehwag 等ICML 2024 · 被引用 15 次
- Descending through a Crowded Valley - Benchmarking Deep Learning OptimizersRobin M. Schmidt, Frank Schneider, Philipp HennigICML 2021 · 被引用 195 次
