SAPIENTML: Synthesizing Machine Learning Pipelines by Learning from Human-Written Solutions
Ripon K. Saha, Akira Ura, Sonal Mahajan, Chenguang Zhu, Linyi Li, Yang Hu, Hiroaki Yoshida, Sarfraz Khurshid, Mukul R. Prasad
摘要
Automatic machine learning, or AutoML, holds the promise of truly democratizing the use of machine learning (ML), by substantially automating the work of data scientists. However, the huge combinatorial search space of candidate pipelines means that current AutoML techniques, generate sub-optimal pipelines, or none at all, especially on large, complex datasets. In this work we propose an AutoML technique SapientML, that can learn from a corpus of existing datasets and their human-written pipelines, and efficiently generate a high-quality pipeline for a predictive task on a new dataset. To combat the search space explosion of AutoML, SapientML employs a novel divide-and-conquer strategy realized as a three-stage program synthesis approach, that reasons on successively smaller search spaces. The first stage uses meta-learning to predict a set of plausible ML components to constitute a pipeline. In the second stage, this is then refined into a small pool of viable concrete pipelines using a pipeline dataflow model derived from the corpus. Dynamically evaluating these few pipelines, in the third stage, provides the best solution. We instantiate SapientML as part of a fully automated tool-chain that creates a cleaned, labeled learning corpus by mining Kaggle, learns from it, and uses the learned models to then synthesize pipelines for new predictive tasks. We have created a training corpus of 1,094 pipelines spanning 170 datasets, and evaluated SapientML on a set of 41 benchmark datasets, including 10 new, large, real-world datasets from Kaggle, and against 3 state-of-the-art AutoML tools and 4 baselines. Our evaluation shows that SapientML produces the best or comparable accuracy on 27 of the benchmarks while the second best tool fails to even produce a pipeline on 9 of the instances. This difference is amplified on the 10 most challenging benchmarks, where Sapi-entML wins on 9 instances with the other tools failing to produce pipelines on 4 or more benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- CoCoNuT: combining context-aware neural translation models using ensemble for program repairThibaud Lutellier, Hung Viet Pham, Lawrence Pang, Yitong Li 等ISSTA 2020 · 被引用 325 次
- Assessing and Restoring Reproducibility of Jupyter NotebooksJiawei Wang, Tzu-yang Kuo, Li Li, Andreas ZellerASE 2020 · 被引用 68 次
- Exploring how deprecated Python library APIs are (not) handledJiawei Wang, Li Li, Kui Liu, Haipeng CaiFSE 2020 · 被引用 52 次
- AutoML Pipeline Selection: Efficiently Navigating the Combinatorial SpaceChengrun Yang, Jicong Fan, Ziyang Wu, Madeleine UdellKDD 2020 · 被引用 29 次
- AMS: generating AutoML search spaces from weak specificationsJosé Pablo Cambronero, Jürgen Cito, Martin C. RinardFSE 2020 · 被引用 12 次
相关 Paper
- Doing More with Less: Characterizing Dataset Downsampling for AutoMLFatjon Zogaj, José Pablo Cambronero, Martin C. Rinard, Jürgen CitoVLDB 2021 · 被引用 20 次
- A Scalable AutoML Approach Based on Graph Neural NetworksMossad Helali, Essam Mansour, Ibrahim Abdelaziz, Julian Dolby 等VLDB 2022 · 被引用 16 次
- DeepLine: AutoML Tool for Pipelines Generation using Deep Reinforcement Learning and Hierarchical Actions FilteringYuval Heffetz, Roman Vainshtein, Gilad Katz, Lior RokachKDD 2020 · 被引用 3 次
- Deep Pipeline Embeddings for AutoMLSebastian Pineda-Arango, Josif GrabockaKDD 2023 · 被引用 6 次
- SubStrat: A Subset-Based Optimization Strategy for Faster AutoMLTeddy Lazebnik, Amit Somech, Abraham Itzhak WeinbergVLDB 2023 · 被引用 23 次
