Selective Data Acquisition in the Wild for Model Charging
Chengliang Chai, Jiabin Liu, Nan Tang, Guoliang Li, Yuyu Luo
Abstract
The lack of sufficient labeled data is a key bottleneck for practitioners in many real-world supervised machine learning (ML) tasks. In this paper, we study a new problem, namely selective data acquisition in the wild for model charging : given a supervised ML task and data in the wild (e.g., enterprise data warehouses, online data repositories, data markets, and so on), the problem is to select labeled data points from the data in the wild as additional train data that can help the ML task. It consists of two steps (Fig. 1). The first step is to discover relevant datasets ( e.g. , tables with similar relational schema), which will result in a set of candidate datasets. Because these candidate datasets come from different sources and may follow different distributions, not all data points they contain can help. The second step is to select which data points from these candidate datasets should be used. We build an end-to-end solution. For step 1, we piggyback off-the-shelf data discovery tools. Technically, our focus is on step 2, for which we propose a solution framework called AutoData. It first clusters all data points from candidate datasets such that each cluster contains similar data points from different sources. It then iteratively picks which cluster to use, samples data points ( i.e. , a mini-batch) from the picked cluster, evaluates the mini-batch, and then revises the search criteria by learning from the feedback ( i.e. , reward) based on the evaluation. We propose a multi-armed bandit based solution and a Deep Q Networks-based reinforcement learning solution. Experiments using both relational and image datasets show the effectiveness of our solutions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers23
- Cost-based or Learning-based? A Hybrid Query Optimizer for Query Plan SelectionXiang Yu, Chengliang Chai, Guoliang Li, Jiabin LiuVLDB 2022 · 82 citations
- HAIChart: Human and AI Paired Visualization SystemYupeng Xie, Yuyu Luo, Guoliang Li, Nan TangVLDB 2024 · 47 citations
- GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete DataChengliang Chai, Jiabin Liu, Nan Tang, Ju Fan et al.SIGMOD 2023 · 37 citations
- LakeBench: A Benchmark for Discovering Joinable and Unionable Tables in Data LakesYuhao Deng, Chengliang Chai, Lei Cao, Qin Yuan et al.VLDB 2024 · 36 citations
- Coresets over Multiple Tables for Feature-rich and Data-efficient Machine LearningJiayi Wang, Chengliang Chai, Nan Tang, Jiabin Liu et al.VLDB 2023 · 31 citations
Builds on7
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
- Data Valuation using Reinforcement LearningJinsung Yoon, Sercan Ömer Arik, Tomas PfisterICML 2020 · 236 citations
- Synthesizing Natural Language to Visualization (NL2VIS) Benchmarks from NL2SQL BenchmarksYuyu Luo, Nan Tang, Guoliang Li, Chengliang Chai et al.SIGMOD 2021 · 90 citations
- Human-in-the-loop Outlier DetectionChengliang Chai, Lei Cao, Guoliang Li, Jian Li et al.SIGMOD 2020 · 57 citations
- Tailoring Data Source Distributions for Fairness-aware Data IntegrationFatemeh Nargesian, Abolfazl Asudeh, H. V. JagadishVLDB 2021 · 51 citations
Related papers
- Data Acquisition for Improving Machine Learning ModelsYifan Li, Xiaohui Yu, Nick KoudasVLDB 2021 · 57 citations
- Mitigating Data Scarcity in Supervised Machine Learning Through Reinforcement Learning Guided Data GenerationChengliang Chai, Kaisen Jin, Nan Tang, Ju Fan et al.ICDE 2024 · 7 citations
- Toward Deep Supervised Anomaly Detection: Reinforcement Learning from Partially Labeled Anomaly DataGuansong Pang, Anton van den Hengel, Chunhua Shen, Longbing CaoKDD 2021 · 90 citations
- Automatically Discovering and Learning New Visual Categories with Ranking StatisticsKai Han, Sylvestre-Alvise Rebuffi, Sébastien Ehrhardt, Andrea Vedaldi et al.ICLR 2020 · 222 citations
- Combining Supervised Learning and Reinforcement Learning for Multi-Label Classification Tasks with Partial LabelsZixia Jia, Junpeng Li, Shichuan Zhang, Anji Liu et al.ACL 2024 · 2 citations
