ARDA: Automatic Relational Data Augmentation for Machine Learning
Nadiia Chepurko, Ryan Marcus, Emanuel Zgraggen, Raul Castro Fernandez, Tim Kraska, David R. Karger
摘要
Automatic machine learning (AML) is a family of techniques to automate the process of training predictive models, aiming to both improve performance and make machine learning more accessible. While many recent works have focused on aspects of the machine learning pipeline like model selection, hyperparameter tuning, and feature selection, relatively few works have focused on automatic data augmentation. Automatic data augmentation involves finding new features relevant to the user's predictive task with minimal "human-in-the-loop" involvement. We present ARDA, an end-to-end system that takes as input a dataset and a data repository, and outputs an augmented data set such that training a predictive model on this augmented dataset results in improved performance. Our system has two distinct components: (1) a framework to search and join data with the input data, based on various attributes of the input, and (2) an efficient feature selection algorithm that prunes out noisy or irrelevant features from the resulting join. We perform an extensive empirical evaluation of different system components and benchmark our feature selection algorithm on real-world datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper36
- Valentine: Evaluating Matching Techniques for Dataset DiscoveryChristos Koutras, George Siachamis, Andra Ionescu, Kyriakos Psarakis 等ICDE 2021 · 被引用 87 次
- Cost-based or Learning-based? A Hybrid Query Optimizer for Query Plan SelectionXiang Yu, Chengliang Chai, Guoliang Li, Jiabin LiuVLDB 2022 · 被引用 82 次
- Efficient Joinable Table Discovery in Data Lakes: A High-Dimensional Similarity-Based ApproachYuyang Dong, Kunihiro Takeoka, Chuan Xiao, Masafumi OyamadaICDE 2021 · 被引用 78 次
- Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and BeyondZhengjie Miao, Yuliang Li, Xiaolan WangSIGMOD 2021 · 被引用 63 次
- Selective Data Acquisition in the Wild for Model ChargingChengliang Chai, Jiabin Liu, Nan Tang, Guoliang Li 等VLDB 2022 · 被引用 62 次
相关 Paper
- Featpilot: Automatic Feature Augmentation on Tabular DataJiaming Liang, Chuan Lei, Xiao Qin, Jiani Zhang 等ICDE 2025
- AutoDS: Towards Human-Centered Automation of Data ScienceDakuo Wang, Josh Andres, Justin D. Weisz, Erick Oduor 等CHI 2021 · 被引用 77 次
- DeepLine: AutoML Tool for Pipelines Generation using Deep Reinforcement Learning and Hierarchical Actions FilteringYuval Heffetz, Roman Vainshtein, Gilad Katz, Lior RokachKDD 2020 · 被引用 3 次
- AutoMMLab: Automatically Generating Deployable Models from Language Instructions for Computer Vision TasksZekang Yang, Wang Zeng, Sheng Jin, Chen Qian 等AAAI 2025 · 被引用 18 次
- An ADMM Based Framework for AutoML Pipeline ConfigurationSijia Liu, Parikshit Ram, Deepak Vijaykeerthy, Djallel Bouneffouf 等AAAI 2020 · 被引用 82 次
