SubStrat: A Subset-Based Optimization Strategy for Faster AutoML
Teddy Lazebnik, Amit Somech, Abraham Itzhak Weinberg
Abstract
Automated machine learning (AutoML) frameworks have become important tools in the data scientist's arsenal, as they dramatically reduce the manual work devoted to the construction of ML pipelines. Such frameworks intelligently search among millions of possible ML pipelines - typically containing feature engineering, model selection, and hyper parameters tuning steps - and finally output an optimal pipeline in terms of predictive accuracy. However, when the dataset is large, each individual configuration takes longer to execute, therefore the overall AutoML running times become increasingly high.
To this end, we present SubStrat, an AutoML optimization strategy that tackles the data size, rather than configuration space. It wraps existing AutoML tools, and instead of executing them directly on the entire dataset, SubStrat uses a genetic-based algorithm to find a small yet representative data subset that preserves a particular characteristic of the full data. It then employs the AutoML tool on the small subset, and finally, it refines the resulting pipeline by executing a restricted, much shorter, AutoML process on the large dataset. Our experimental results, performed on three popular AutoML frameworks, Auto-Sklearn, TPOT, and H2O show that SubStrat reduces their running times by 76.3% (on average), with only a 4.15% average decrease in the accuracy of the resulting ML pipeline.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4448625d-fb88-4ab6-9739-3f65ed7f087fCited by top-tier papers3
- Datamap-Driven Tabular Coreset Selection for Classifier TrainingAviv Hadar, Tova Milo, Kathy RazmadzeVLDB 2025 · 6 citations
- ADELA: Accelerating Evolutionary Design of Machine Learning Pipelines with the Accompanying Surrogate ModelYang Gu, Jian Cao, Hengyu You, Nengjun Zhu et al.AAAI 2025 · 1 citation
- CAPS: Cost-Aware ML Pipeline SelectionAntonios Kontaxakis, Dimitris Sacharidis, Alberto Abelló, Sergi Nadal et al.VLDB 2026
Builds on3
- AutoML-Zero: Evolving Machine Learning Algorithms From ScratchEsteban Real, Chen Liang, David R. So, Quoc V. LeICML 2020 · 265 citations
- Frugal Optimization for Cost-related HyperparametersQingyun Wu, Chi Wang, Silu HuangAAAI 2021 · 51 citations
- DeepLine: AutoML Tool for Pipelines Generation using Deep Reinforcement Learning and Hierarchical Actions FilteringYuval Heffetz, Roman Vainshtein, Gilad Katz, Lior RokachKDD 2020 · 3 citations
Related papers
- Doing More with Less: Characterizing Dataset Downsampling for AutoMLFatjon Zogaj, José Pablo Cambronero, Martin C. Rinard, Jürgen CitoVLDB 2021 · 20 citations
- SAPIENTML: Synthesizing Machine Learning Pipelines by Learning from Human-Written SolutionsRipon K. Saha, Akira Ura, Sonal Mahajan, Chenguang Zhu et al.ICSE 2022 · 11 citations
- VolcanoML: Speeding up End-to-End AutoML via Scalable Search Space DecompositionYang Li, Yu Shen, Wentao Zhang, Jiawei Jiang et al.VLDB 2021 · 55 citations
- AUTOMATA: Gradient Based Data Subset Selection for Compute-Efficient Hyper-parameter TuningKrishnaTeja Killamsetty, Guttu Sai Abhishek, Aakriti, Ganesh Ramakrishnan et al.NeurIPS 2022 · 37 citations
- DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular DataPeng Li, Zhiyi Chen, Xu Chu, Kexin RongSIGMOD 2023 · 24 citations
