Optimizing Data Acquisition to Enhance Machine Learning Performance
Tingting Wang, Shixun Huang, Zhifeng Bao, J. Shane Culpepper, Volkan Dedeoglu, Reza Arablouei
Abstract
In this paper, we study how to acquire labeled data points from a large data pool to enrich a training set for enhancing supervised machine learning (ML) performance. The state-of-the-art solution is the clustering-based training set selection (CTS) algorithm, which initially clusters the data points in a data pool and subsequently selects new data points from clusters. The efficiency of CTS is constrained by its frequent retraining of the target ML model, and the effectiveness is limited by the selection criteria, which represent the state of data points within each cluster and impose a restriction of selecting only one cluster in each iteration. To overcome these limitations, we propose a new algorithm, called CTS with incremental estimation of adaptive score (IAS). IAS employs online learning, enabling incremental model updates by using new data, and eliminating the need to fully retrain the target model, and hence improves the efficiency. To enhance the effectiveness of IAS, we introduce adaptive score estimation, which serves as novel selection criteria to identify clusters and select new data points by balancing trade-offs between exploitation and exploration during data acquisition. To further enhance the effectiveness of IAS, we introduce a new adaptive mini-batch selection method that, in each iteration, selects data points from multiple clusters rather than a single cluster, hence eliminating the potential bias due to using only one cluster. By integrating this method into the IAS algorithm, we propose a novel algorithm termed IAS with adaptive mini-batch selection (IAS-AMS). Experimental results highlight the superior effectiveness of IAS-AMS, with IAS also outperforming other competing algorithms. In terms of efficiency, IAS takes the lead, while the efficiency of IAS-AMS is on par with that of the existing CTS algorithm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85644eea-ab4a-4e48-972f-4e0d003dea62Cited by top-tier papers5
- Distinctiveness Maximization in Datasets AssemblageTingting Wang, Shixun Huang, Zhifeng Bao, J. Shane Culpepper et al.WWW 2025 · 3 citations
- A Cost-Effective LLM-based Approach to Identify Wildlife Trafficking in Online MarketplacesJuliana Silva Barbosa, Ulhas Gondhali, Gohar Petrossian, Kinshuk Sharma et al.SIGMOD 2025 · 3 citations
- Analyzing and Optimizing Perturbation of DP-SGD GeometricallyJiawei Duan, Haibo Hu, Qingqing Ye, Xinyue SunICDE 2025 · 3 citations
- Shapley Value Estimation based on Differential MatrixJunyuan Pang, Jian Pei, Haocheng Xia, Xiang Li et al.SIGMOD 2025 · 2 citations
- Data-efficient Online Training for Direct Alignment in LLMsChi Zhang, Jiacheng Wang, Kun He, Chengliang Chai et al.VLDB 2026
Builds on17
- Data Valuation using Reinforcement LearningJinsung Yoon, Sercan Ömer Arik, Tomas PfisterICML 2020 · 236 citations
- COSTA: Covariance-Preserving Feature Augmentation for Graph Contrastive LearningYifei Zhang, Hao Zhu, Zixing Song, Piotr Koniusz et al.KDD 2022 · 95 citations
- Synthesizing Natural Language to Visualization (NL2VIS) Benchmarks from NL2SQL BenchmarksYuyu Luo, Nan Tang, Guoliang Li, Chengliang Chai et al.SIGMOD 2021 · 90 citations
- Selective Data Acquisition in the Wild for Model ChargingChengliang Chai, Jiabin Liu, Nan Tang, Guoliang Li et al.VLDB 2022 · 62 citations
- Towards Sustainable Learning: Coresets for Data-efficient Deep LearningYu Yang, Hao Kang, Baharan MirzasoleimanICML 2023 · 58 citations
Related papers
- Class Balance Matters to Active Class-Incremental LearningZitong Huang, Ze Chen, Yuanze Li, Bowen Dong et al.ACM MM 2024 · 8 citations
- Active Learning from the WebRyoma SatoWWW 2023
- Making Look-Ahead Active Learning Strategies Feasible with Neural Tangent KernelsMohamad Amin Mohamadi, Wonho Bae, Danica J. SutherlandNeurIPS 2022 · 32 citations
- Algorithm Selection for Deep Active Learning with Imbalanced DatasetsJifan Zhang, Shuai Shao, Saurabh Verma, Robert D. NowakNeurIPS 2023 · 38 citations
- A Neural Pre-Conditioning Active Learning Algorithm to Reduce Label ComplexitySeo Taek Kong, Soomin Jeon, Dongbin Na, Jaewon Lee et al.NeurIPS 2022 · 7 citations
