Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning
Jiayi Wang, Chengliang Chai, Nan Tang, Jiabin Liu, Guoliang Li
Abstract
Successful machine learning (ML) needs to learn from good data. However, one common issue about train data for ML practitioners is the lack of good features. To mitigate this problem, feature augmentation is often employed by joining with (or enriching features from) multiple tables, so as to become feature-rich ML. A consequent problem is that the enriched train data may contain too many tuples, especially if the feature augmentation is obtained through 1 (or many)-to-many or fuzzy joins. Training an ML model with a very large train dataset is data-inefficient. Coreset is often used to achieve data-efficient ML training, which selects a small subset of train data that can theoretically and practically perform similarly as using the full dataset. However, coreset selection over a large train dataset is also known to be time-consuming. In this paper, we aim at achieving both feature-rich ML through feature augmentation and data-efficient ML through coreset selection. In order to avoid time-consuming coreset selection over a feature augmented (or fully materialized) table, we propose to efficiently select the coreset without materializing the augmented table. Note that coreset selection typically uses weighted gradients of the subset to approximate the full gradient of the entire train dataset. Our key idea is that the gradient computation for coreset selection of the augmented table can be pushed down to partial feature similarity of tuples within each individual table, without join materialization. These partial feature similarity values can be aggregated to estimate the gradient of the augmented table, which is upper bounded with provable theoretical guarantees. Extensive experiments show that our method can improve the efficiency by nearly 2 orders of magnitudes, while keeping almost the same accuracy as training with the fully augmented train data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ad17f5b7-4322-44bf-952f-6530d347ab08Cited by top-tier papers11
- Cost-based or Learning-based? A Hybrid Query Optimizer for Query Plan SelectionXiang Yu, Chengliang Chai, Guoliang Li, Jiabin LiuVLDB 2022 · 82 citations
- GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete DataChengliang Chai, Jiabin Liu, Nan Tang, Ju Fan et al.SIGMOD 2023 · 37 citations
- LakeBench: A Benchmark for Discovering Joinable and Unionable Tables in Data LakesYuhao Deng, Chengliang Chai, Lei Cao, Qin Yuan et al.VLDB 2024 · 36 citations
- Optimizing Data Acquisition to Enhance Machine Learning PerformanceTingting Wang, Shixun Huang, Zhifeng Bao, J. Shane Culpepper et al.VLDB 2024 · 13 citations
- MisDetect: Iterative Mislabel Detection using Early LossYuhao Deng, Chengliang Chai, Lei Cao, Nan Tang et al.VLDB 2024 · 13 citations
Builds on9
- Bao: Making Learned Query Optimization PracticalRyan Marcus, Parimarjan Negi, Hongzi Mao, Nesime Tatbul et al.SIGMOD 2021 · 242 citations
- FairBatch: Batch Selection for Model FairnessYuji Roh, Kangwook Lee, Steven Euijong Whang, Changho SuhICLR 2021 · 156 citations
- NeuroCard: One Cardinality Estimator for All TablesZongheng Yang, Amog Kamsetty, Sifei Luan, Eric Liang et al.VLDB 2021 · 138 citations
- Coresets for Robust Training of Deep Neural Networks against Noisy LabelsBaharan Mirzasoleiman, Kaidi Cao, Jure LeskovecNeurIPS 2020 · 99 citations
- Selective Data Acquisition in the Wild for Model ChargingChengliang Chai, Jiabin Liu, Nan Tang, Guoliang Li et al.VLDB 2022 · 62 citations
Related papers
- Efficient Coreset Selection with Cluster-based MethodsChengliang Chai, Jiayi Wang, Nan Tang, Ye Yuan et al.KDD 2023 · 19 citations
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
- Coresets for Relational Data and The ApplicationsJiaxiang Chen, Qingyuan Yang, Ruomin Huang, Hu DingNeurIPS 2022 · 10 citations
- A Novel Sequential Coreset Method for Gradient Descent AlgorithmsJiawei Huang, Ruomin Huang, Wenjie Liu, Nikolaos M. Freris et al.ICML 2021 · 20 citations
- Adaptive Second Order Coresets for Data-efficient Machine LearningOmead Pooladzandi, David Davini, Baharan MirzasoleimanICML 2022 · 83 citations
