SubStrat: A Subset-Based Optimization Strategy for Faster AutoML
Teddy Lazebnik, Amit Somech, Abraham Itzhak Weinberg
摘要
Automated machine learning (AutoML) frameworks have become important tools in the data scientist's arsenal, as they dramatically reduce the manual work devoted to the construction of ML pipelines. Such frameworks intelligently search among millions of possible ML pipelines - typically containing feature engineering, model selection, and hyper parameters tuning steps - and finally output an optimal pipeline in terms of predictive accuracy. However, when the dataset is large, each individual configuration takes longer to execute, therefore the overall AutoML running times become increasingly high.
To this end, we present SubStrat, an AutoML optimization strategy that tackles the data size, rather than configuration space. It wraps existing AutoML tools, and instead of executing them directly on the entire dataset, SubStrat uses a genetic-based algorithm to find a small yet representative data subset that preserves a particular characteristic of the full data. It then employs the AutoML tool on the small subset, and finally, it refines the resulting pipeline by executing a restricted, much shorter, AutoML process on the large dataset. Our experimental results, performed on three popular AutoML frameworks, Auto-Sklearn, TPOT, and H2O show that SubStrat reduces their running times by 76.3% (on average), with only a 4.15% average decrease in the accuracy of the resulting ML pipeline.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Datamap-Driven Tabular Coreset Selection for Classifier TrainingAviv Hadar, Tova Milo, Kathy RazmadzeVLDB 2025 · 被引用 6 次
- ADELA: Accelerating Evolutionary Design of Machine Learning Pipelines with the Accompanying Surrogate ModelYang Gu, Jian Cao, Hengyu You, Nengjun Zhu 等AAAI 2025 · 被引用 1 次
- CAPS: Cost-Aware ML Pipeline SelectionAntonios Kontaxakis, Dimitris Sacharidis, Alberto Abelló, Sergi Nadal 等VLDB 2026
它引用的顶会 Paper3
- AutoML-Zero: Evolving Machine Learning Algorithms From ScratchEsteban Real, Chen Liang, David R. So, Quoc V. LeICML 2020 · 被引用 265 次
- Frugal Optimization for Cost-related HyperparametersQingyun Wu, Chi Wang, Silu HuangAAAI 2021 · 被引用 51 次
- DeepLine: AutoML Tool for Pipelines Generation using Deep Reinforcement Learning and Hierarchical Actions FilteringYuval Heffetz, Roman Vainshtein, Gilad Katz, Lior RokachKDD 2020 · 被引用 3 次
相关 Paper
- Doing More with Less: Characterizing Dataset Downsampling for AutoMLFatjon Zogaj, José Pablo Cambronero, Martin C. Rinard, Jürgen CitoVLDB 2021 · 被引用 20 次
- SAPIENTML: Synthesizing Machine Learning Pipelines by Learning from Human-Written SolutionsRipon K. Saha, Akira Ura, Sonal Mahajan, Chenguang Zhu 等ICSE 2022 · 被引用 11 次
- VolcanoML: Speeding up End-to-End AutoML via Scalable Search Space DecompositionYang Li, Yu Shen, Wentao Zhang, Jiawei Jiang 等VLDB 2021 · 被引用 55 次
- AUTOMATA: Gradient Based Data Subset Selection for Compute-Efficient Hyper-parameter TuningKrishnaTeja Killamsetty, Guttu Sai Abhishek, Aakriti, Ganesh Ramakrishnan 等NeurIPS 2022 · 被引用 37 次
- DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular DataPeng Li, Zhiyi Chen, Xu Chu, Kexin RongSIGMOD 2023 · 被引用 24 次
