AMS: generating AutoML search spaces from weak specifications
José Pablo Cambronero, Jürgen Cito, Martin C. Rinard
摘要
We consider a usage model for automated machine learning (Au-toML) in which users can influence the generated pipeline by providing a weak pipeline specification: an unordered set of API components from which the AutoML system draws the components it places into the generated pipeline. Such specifications allow users to express preferences over the components that appear in the pipeline, for example a desire for interpretable components to appear in the pipeline. We present AMS, an approach to automatically strengthen weak specifications to include unspecified complementary and functionally related API components, populate the space of hyperparameters and their values, and pair this configuration with a search procedure to produce a strong pipeline specification: a full description of the search space for candidate pipelines. AMS uses normalized pointwise mutual information on a code corpus to identify complementary components, BM25 as a lexical similarity score over the target API's documentation to identify functionally related components, and frequency distributions in the code corpus to extract key hyperparameters and values. We show that strengthened specifications can produce pipelines that outperform the pipelines generated from the initial weak specification and an expert-annotated variant, while producing pipelines that still reflect the user preferences captured in the original weak specification.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Doing More with Less: Characterizing Dataset Downsampling for AutoMLFatjon Zogaj, José Pablo Cambronero, Martin C. Rinard, Jürgen CitoVLDB 2021 · 被引用 20 次
- Resource-Guided Configuration Space Reduction for Deep Learning ModelsYanjie Gao, Yonghao Zhu, Hongyu Zhang, Haoxiang Lin 等ICSE 2021 · 被引用 17 次
- Manas: Mining Software Repositories to Assist AutoMLGiang Nguyen, Md Johirul Islam, Rangeet Pan, Hridesh RajanICSE 2022 · 被引用 15 次
- Fix Fairness, Don't Ruin Accuracy: Performance Aware Fairness Repair using AutoMLGiang Nguyen, Sumon Biswas, Hridesh RajanFSE 2023 · 被引用 15 次
- SAPIENTML: Synthesizing Machine Learning Pipelines by Learning from Human-Written SolutionsRipon K. Saha, Akira Ura, Sonal Mahajan, Chenguang Zhu 等ICSE 2022 · 被引用 11 次
相关 Paper
- AutoML-Agent: A Multi-Agent LLM Framework for Full-Pipeline AutoMLPatara Trirat, Wonyong Jeong, Sung Ju HwangICML 2025
- Automating API Documentation from Crowdsourced KnowledgeBonan Kou, Zijie Zhou, Muhao Chen, Tianyi ZhangICSE 2026
- Pipeline Combinators for Gradual AutoMLGuillaume Baudart, Martin Hirzel, Kiran Kate, Parikshit Ram 等NeurIPS 2021 · 被引用 27 次
- Improving API Knowledge Discovery with ML: A Case Study of Comparable API MethodsDaye Nam, Brad A. Myers, Bogdan Vasilescu, Vincent J. HellendoornICSE 2023 · 被引用 6 次
- Whither AutoML? Understanding the Role of Automation in Machine Learning WorkflowsDoris Xin, Eva Yiwei Wu, Doris Jung Lin Lee, Niloufar Salehi 等CHI 2021 · 被引用 103 次
