ADELA: Accelerating Evolutionary Design of Machine Learning Pipelines with the Accompanying Surrogate Model
Yang Gu, Jian Cao, Hengyu You, Nengjun Zhu, Shiyou Qian
Abstract
The end-to-end automated design of machine learning (ML) pipelines significantly reduces the workload for data scientists and democratizes ML for non-experts. Evolutionary algorithm (EA)-based automated ML (AutoML) systems, a prominent category of AutoML, often face inefficiencies due to the costly fitness evaluation of candidate ML pipelines. Although surrogate models have been employed to approximate the true performance of pipelines more quickly, a key challenge remains in effectively bridging the semantic gap between the heterogeneous features of datasets and pipelines. To address this issue, we propose ADELA, a novel accompanying surrogate-based optimization strategy that accelerates EA-based AutoML while retaining the performance of the resulting pipelines. ADELA operates in two phases: Offline, leveraging a high-quality curated pipeline corpus to meta-learn an accompanying surrogate model; and Online, selecting the accompanying pipeline and using the learned model to predict the performance of evaluation pipelines instead of executing them. The accompanying mechanism effectively mitigates the semantic gap between datasets and pipelines, enabling ADELA to reduce computation times by an average of 73.66% while retaining 98.78% of the final pipeline performance, as demonstrated in extensive experimental evaluations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e450ff5f-fd8f-4c1b-a2e9-e217324beae3Builds on9
- Transformers Can Do Bayesian InferenceSamuel Müller, Noah Hollmann, Sebastian Pineda-Arango, Josif Grabocka et al.ICLR 2022 · 287 citations
- Large Language Models of Code Fail at Completing Code with Potential BugsTuan Dinh, Jinman Zhao, Samson Tan, Renato Negrinho et al.NeurIPS 2023 · 59 citations
- VolcanoML: Speeding up End-to-End AutoML via Scalable Search Space DecompositionYang Li, Yu Shen, Wentao Zhang, Jiawei Jiang et al.VLDB 2021 · 55 citations
- Serving and Optimizing Machine Learning Workflows on Heterogeneous InfrastructuresYongji Wu, Matthew Lentz, Danyang Zhuo, Yao LuVLDB 2023 · 31 citations
- SubStrat: A Subset-Based Optimization Strategy for Faster AutoMLTeddy Lazebnik, Amit Somech, Abraham Itzhak WeinbergVLDB 2023 · 23 citations
Related papers
- SAPIENTML: Synthesizing Machine Learning Pipelines by Learning from Human-Written SolutionsRipon K. Saha, Akira Ura, Sonal Mahajan, Chenguang Zhu et al.ICSE 2022 · 11 citations
- Zero-shot AutoML with Pretrained ModelsEkrem Öztürk, Fabio Ferreira, Hadi S. Jomaa, Lars Schmidt-Thieme et al.ICML 2022 · 17 citations
- Deep Pipeline Embeddings for AutoMLSebastian Pineda-Arango, Josif GrabockaKDD 2023 · 6 citations
- CAPS: Cost-Aware ML Pipeline SelectionAntonios Kontaxakis, Dimitris Sacharidis, Alberto Abelló, Sergi Nadal et al.VLDB 2026
- DeepLine: AutoML Tool for Pipelines Generation using Deep Reinforcement Learning and Hierarchical Actions FilteringYuval Heffetz, Roman Vainshtein, Gilad Katz, Lior RokachKDD 2020 · 3 citations
