Learning meta-features for AutoML
Herilalaina Rakotoarison, Louisot Milijaona, Andry Rasoanaivo, Michèle Sebag, Marc Schoenauer
Abstract
This paper tackles the AutoML problem, aimed to automatically select an ML algorithm and its hyper-parameter configuration most appropriate to the dataset at hand. The proposed approach, MetaBu, learns new meta-features via an Optimal Transport procedure, aligning the manually designed meta-features with the space of distributions on the hyper-parameter configurations. MetaBu meta-features, learned once and for all, induce a topology on the set of datasets that is exploited to define a distribution of promising hyper-parameter configurations amenable to AutoML. Experiments on the OpenML CC-18 benchmark demonstrate that using MetaBu meta-features boosts the performance of state of the art AutoML systems, (Feurer et al. 2015) and Probabilistic Matrix Factorization (Fusi et al. 2018). Furthermore, the inspection of MetaBu meta-features gives some hints into when an ML algorithm does well. Finally, the topology based on MetaBu meta-features enables to estimate the intrinsic dimensionality of the OpenML benchmark w.r.t. a given ML algorithm or pipeline.
- equal contribution 1 An ML pipeline consists of a data preparation stage followed by the model learning stage. Each stage involves a number of options and a varying number of hyper-parameters, depending on the former selected options. Terms ML pipeline and ML algorithm will be used interchangeably in the remainder of the paper.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c7170220-93ec-46f3-9a8d-631e1b120087Cited by top-tier papers3
- Deep Ranking Ensembles for Hyperparameter OptimizationAbdus Salam Khazi, Sebastian Pineda-Arango, Josif GrabockaICLR 2023 · 1 citation
- ADELA: Accelerating Evolutionary Design of Machine Learning Pipelines with the Accompanying Surrogate ModelYang Gu, Jian Cao, Hengyu You, Nengjun Zhu et al.AAAI 2025 · 1 citation
- On the Hyperparameter Loss Landscapes of Machine Learning Models: An Exploratory StudyMingyu Huang, Ke LiKDD 2025
Builds on3
- Geometric Dataset Distances via Optimal TransportDavid Alvarez-Melis, Nicolò FusiNeurIPS 2020 · 267 citations
- On Learning Sets of Symmetric ElementsHaggai Maron, Or Litany, Gal Chechik, Ethan FetayaICML 2020 · 148 citations
- Improving Relational Regularized Autoencoders with Spherical Sliced Fused Gromov WassersteinKhai Nguyen, Son Nguyen, Nhat Ho, Tung Pham et al.ICLR 2021 · 21 citations
Related papers
- A Scalable AutoML Approach Based on Graph Neural NetworksMossad Helali, Essam Mansour, Ibrahim Abdelaziz, Julian Dolby et al.VLDB 2022 · 16 citations
- SAPIENTML: Synthesizing Machine Learning Pipelines by Learning from Human-Written SolutionsRipon K. Saha, Akira Ura, Sonal Mahajan, Chenguang Zhu et al.ICSE 2022 · 11 citations
- Searching for Machine Learning Pipelines Using a Context-Free GrammarRadu Marinescu, Akihiro Kishimoto, Parikshit Ram, Ambrish Rawat et al.AAAI 2021 · 18 citations
- Deep Pipeline Embeddings for AutoMLSebastian Pineda-Arango, Josif GrabockaKDD 2023 · 6 citations
- An ADMM Based Framework for AutoML Pipeline ConfigurationSijia Liu, Parikshit Ram, Deepak Vijaykeerthy, Djallel Bouneffouf et al.AAAI 2020 · 82 citations
