GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete Data
Chengliang Chai, Jiabin Liu, Nan Tang, Ju Fan, Dongjing Miao, Jiayi Wang, Yuyu Luo, Guoliang Li
摘要
Given a dataset with incomplete data (e.g., missing values), training a machine learning model over the incomplete data requires two steps. First, it requires a data-effective step that cleans the data in order to improve the data quality (and the model quality on the cleaned data). Second, it requires a data-efficient step that selects a core subset of the data (called coreset) such that the trained models on the entire data and the coreset have similar model quality, in order to improve the training efficiency. The first-data-effective-then-data-efficient methods are too costly, because they are expensive to clean the whole data; while the first-data-efficient-then-data-effective methods have low model quality, because they cannot select high-quality coreset for incomplete data. In this paper, we investigate the problem of coreset selection over incomplete data for data-effective and data-efficient machine learning. The essential challenge is how to model the incomplete data for selecting high-quality coreset. To this end, we propose the GoodCore framework towards selecting a good coreset over incomplete data with low cost. To model the unknown complete data, we utilize the combinations of possible repairs as possible worlds of the incomplete data. Based on possible worlds, GoodCore selects an expected optimal coreset through gradient approximation without training ML models. We formally define the expected optimal coreset selection problem, prove its NP-hardness, and propose a greedy algorithm with an approximation ratio. To make GoodCore more efficient, we further propose optimization methods that incorporate human-in-the-loop imputation or automatic imputation method into our framework. Experimental results show the effectiveness and efficiency of our framework with low cost.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- LakeBench: A Benchmark for Discovering Joinable and Unionable Tables in Data LakesYuhao Deng, Chengliang Chai, Lei Cao, Qin Yuan 等VLDB 2024 · 被引用 36 次
- LEAD: Iterative Data Selection for Efficient LLM Instruction TuningXiaotian Lin, Yanlin Qi, Yizhang Zhu, Themis Palpanas 等VLDB 2026 · 被引用 16 次
- Optimizing Data Acquisition to Enhance Machine Learning PerformanceTingting Wang, Shixun Huang, Zhifeng Bao, J. Shane Culpepper 等VLDB 2024 · 被引用 13 次
- MisDetect: Iterative Mislabel Detection using Early LossYuhao Deng, Chengliang Chai, Lei Cao, Nan Tang 等VLDB 2024 · 被引用 13 次
- QCore: Data-Efficient, On-Device Continual Calibration for Quantized ModelsDavid Campos, Bin Yang, Tung Kieu, Miao Zhang 等VLDB 2024 · 被引用 11 次
它引用的顶会 Paper16
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 被引用 494 次
- GRAD-MATCH: Gradient Matching based Data Subset Selection for Efficient Deep Model TrainingKrishnaTeja Killamsetty, Durga Sivasubramanian, Ganesh Ramakrishnan, Abir De 等ICML 2021 · 被引用 305 次
- GLISTER: Generalization based Data Subset Selection for Efficient and Robust LearningKrishnaTeja Killamsetty, Durga Sivasubramanian, Ganesh Ramakrishnan, Rishabh K. IyerAAAI 2021 · 被引用 300 次
- Natural Language to Visualization by Neural Machine TranslationYuyu Luo, Nan Tang, Guoliang Li, Jiawei Tang 等IEEE VIS 2021 · 被引用 145 次
- CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification TasksPeng Li, Xi Rao, Jennifer Blase, Yue Zhang 等ICDE 2021 · 被引用 127 次
相关 Paper
- Efficient Coreset Selection with Cluster-based MethodsChengliang Chai, Jiayi Wang, Nan Tang, Ye Yuan 等KDD 2023 · 被引用 19 次
- Coresets over Multiple Tables for Feature-rich and Data-efficient Machine LearningJiayi Wang, Chengliang Chai, Nan Tang, Jiabin Liu 等VLDB 2023 · 被引用 31 次
- Certain and Approximately Certain Models for Statistical LearningCheng Zhen, Nischal Aryal, Arash Termehchy, Amandeep Singh ChabadaSIGMOD 2024 · 被引用 4 次
- RETRIEVE: Coreset Selection for Efficient and Robust Semi-Supervised LearningKrishnaTeja Killamsetty, Xujiang Zhao, Feng Chen, Rishabh K. IyerNeurIPS 2021 · 被引用 115 次
- Coresets for Relational Data and The ApplicationsJiaxiang Chen, Qingyuan Yang, Ruomin Huang, Hu DingNeurIPS 2022 · 被引用 10 次
