Optimizing Data Acquisition to Enhance Machine Learning Performance
Tingting Wang, Shixun Huang, Zhifeng Bao, J. Shane Culpepper, Volkan Dedeoglu, Reza Arablouei
摘要
In this paper, we study how to acquire labeled data points from a large data pool to enrich a training set for enhancing supervised machine learning (ML) performance. The state-of-the-art solution is the clustering-based training set selection (CTS) algorithm, which initially clusters the data points in a data pool and subsequently selects new data points from clusters. The efficiency of CTS is constrained by its frequent retraining of the target ML model, and the effectiveness is limited by the selection criteria, which represent the state of data points within each cluster and impose a restriction of selecting only one cluster in each iteration. To overcome these limitations, we propose a new algorithm, called CTS with incremental estimation of adaptive score (IAS). IAS employs online learning, enabling incremental model updates by using new data, and eliminating the need to fully retrain the target model, and hence improves the efficiency. To enhance the effectiveness of IAS, we introduce adaptive score estimation, which serves as novel selection criteria to identify clusters and select new data points by balancing trade-offs between exploitation and exploration during data acquisition. To further enhance the effectiveness of IAS, we introduce a new adaptive mini-batch selection method that, in each iteration, selects data points from multiple clusters rather than a single cluster, hence eliminating the potential bias due to using only one cluster. By integrating this method into the IAS algorithm, we propose a novel algorithm termed IAS with adaptive mini-batch selection (IAS-AMS). Experimental results highlight the superior effectiveness of IAS-AMS, with IAS also outperforming other competing algorithms. In terms of efficiency, IAS takes the lead, while the efficiency of IAS-AMS is on par with that of the existing CTS algorithm.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Distinctiveness Maximization in Datasets AssemblageTingting Wang, Shixun Huang, Zhifeng Bao, J. Shane Culpepper 等WWW 2025 · 被引用 3 次
- A Cost-Effective LLM-based Approach to Identify Wildlife Trafficking in Online MarketplacesJuliana Silva Barbosa, Ulhas Gondhali, Gohar Petrossian, Kinshuk Sharma 等SIGMOD 2025 · 被引用 3 次
- Analyzing and Optimizing Perturbation of DP-SGD GeometricallyJiawei Duan, Haibo Hu, Qingqing Ye, Xinyue SunICDE 2025 · 被引用 3 次
- Shapley Value Estimation based on Differential MatrixJunyuan Pang, Jian Pei, Haocheng Xia, Xiang Li 等SIGMOD 2025 · 被引用 2 次
- Data-efficient Online Training for Direct Alignment in LLMsChi Zhang, Jiacheng Wang, Kun He, Chengliang Chai 等VLDB 2026
它引用的顶会 Paper17
- Data Valuation using Reinforcement LearningJinsung Yoon, Sercan Ömer Arik, Tomas PfisterICML 2020 · 被引用 236 次
- COSTA: Covariance-Preserving Feature Augmentation for Graph Contrastive LearningYifei Zhang, Hao Zhu, Zixing Song, Piotr Koniusz 等KDD 2022 · 被引用 95 次
- Synthesizing Natural Language to Visualization (NL2VIS) Benchmarks from NL2SQL BenchmarksYuyu Luo, Nan Tang, Guoliang Li, Chengliang Chai 等SIGMOD 2021 · 被引用 90 次
- Selective Data Acquisition in the Wild for Model ChargingChengliang Chai, Jiabin Liu, Nan Tang, Guoliang Li 等VLDB 2022 · 被引用 62 次
- Towards Sustainable Learning: Coresets for Data-efficient Deep LearningYu Yang, Hao Kang, Baharan MirzasoleimanICML 2023 · 被引用 58 次
相关 Paper
- Class Balance Matters to Active Class-Incremental LearningZitong Huang, Ze Chen, Yuanze Li, Bowen Dong 等ACM MM 2024 · 被引用 8 次
- Active Learning from the WebRyoma SatoWWW 2023
- Making Look-Ahead Active Learning Strategies Feasible with Neural Tangent KernelsMohamad Amin Mohamadi, Wonho Bae, Danica J. SutherlandNeurIPS 2022 · 被引用 32 次
- Algorithm Selection for Deep Active Learning with Imbalanced DatasetsJifan Zhang, Shuai Shao, Saurabh Verma, Robert D. NowakNeurIPS 2023 · 被引用 38 次
- A Neural Pre-Conditioning Active Learning Algorithm to Reduce Label ComplexitySeo Taek Kong, Soomin Jeon, Dongbin Na, Jaewon Lee 等NeurIPS 2022 · 被引用 7 次
