Learning meta-features for AutoML
Herilalaina Rakotoarison, Louisot Milijaona, Andry Rasoanaivo, Michèle Sebag, Marc Schoenauer
摘要
This paper tackles the AutoML problem, aimed to automatically select an ML algorithm and its hyper-parameter configuration most appropriate to the dataset at hand. The proposed approach, MetaBu, learns new meta-features via an Optimal Transport procedure, aligning the manually designed meta-features with the space of distributions on the hyper-parameter configurations. MetaBu meta-features, learned once and for all, induce a topology on the set of datasets that is exploited to define a distribution of promising hyper-parameter configurations amenable to AutoML. Experiments on the OpenML CC-18 benchmark demonstrate that using MetaBu meta-features boosts the performance of state of the art AutoML systems, (Feurer et al. 2015) and Probabilistic Matrix Factorization (Fusi et al. 2018). Furthermore, the inspection of MetaBu meta-features gives some hints into when an ML algorithm does well. Finally, the topology based on MetaBu meta-features enables to estimate the intrinsic dimensionality of the OpenML benchmark w.r.t. a given ML algorithm or pipeline.
- equal contribution 1 An ML pipeline consists of a data preparation stage followed by the model learning stage. Each stage involves a number of options and a varying number of hyper-parameters, depending on the former selected options. Terms ML pipeline and ML algorithm will be used interchangeably in the remainder of the paper.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Deep Ranking Ensembles for Hyperparameter OptimizationAbdus Salam Khazi, Sebastian Pineda-Arango, Josif GrabockaICLR 2023 · 被引用 1 次
- ADELA: Accelerating Evolutionary Design of Machine Learning Pipelines with the Accompanying Surrogate ModelYang Gu, Jian Cao, Hengyu You, Nengjun Zhu 等AAAI 2025 · 被引用 1 次
- On the Hyperparameter Loss Landscapes of Machine Learning Models: An Exploratory StudyMingyu Huang, Ke LiKDD 2025
它引用的顶会 Paper3
- Geometric Dataset Distances via Optimal TransportDavid Alvarez-Melis, Nicolò FusiNeurIPS 2020 · 被引用 267 次
- On Learning Sets of Symmetric ElementsHaggai Maron, Or Litany, Gal Chechik, Ethan FetayaICML 2020 · 被引用 148 次
- Improving Relational Regularized Autoencoders with Spherical Sliced Fused Gromov WassersteinKhai Nguyen, Son Nguyen, Nhat Ho, Tung Pham 等ICLR 2021 · 被引用 21 次
相关 Paper
- A Scalable AutoML Approach Based on Graph Neural NetworksMossad Helali, Essam Mansour, Ibrahim Abdelaziz, Julian Dolby 等VLDB 2022 · 被引用 16 次
- SAPIENTML: Synthesizing Machine Learning Pipelines by Learning from Human-Written SolutionsRipon K. Saha, Akira Ura, Sonal Mahajan, Chenguang Zhu 等ICSE 2022 · 被引用 11 次
- Searching for Machine Learning Pipelines Using a Context-Free GrammarRadu Marinescu, Akihiro Kishimoto, Parikshit Ram, Ambrish Rawat 等AAAI 2021 · 被引用 18 次
- Deep Pipeline Embeddings for AutoMLSebastian Pineda-Arango, Josif GrabockaKDD 2023 · 被引用 6 次
- An ADMM Based Framework for AutoML Pipeline ConfigurationSijia Liu, Parikshit Ram, Deepak Vijaykeerthy, Djallel Bouneffouf 等AAAI 2020 · 被引用 82 次
