Data Acquisition for Improving Model Confidence
Yifan Li, Xiaohui Yu, Nick Koudas
Abstract
In recent years, there has been a growing recognition that high-quality training data is crucial for the performance of machine learning models. This awareness has catalyzed both research endeavors and industrial initiatives dedicated to data acquisition to enhance diverse dimensions of model performance. Among these dimensions, model confidence holds paramount importance; however, it has often been overlooked in prior investigations into data acquisition methodologies. To address this gap, our work focuses on improving the data acquisition process with the goal of enhancing the confidence of Machine Learning models. Specifically, we operate within a practical context where limited samples can be obtained from a large data pool. We employ well-established model confidence metrics as our foundation, and we propose two methodologies, Bulk Acquisition (BA) and Sequential Acquisition (SA), each geared towards identifying the sets of samples that yield the most substantial gains in model confidence. Recognizing the complexity of BA and SA, we introduce two efficient approximate methods, namely kNN-BA and kNN-SA, restricting data acquisition to promising subsets within the data pool. To broaden the applicability of our solutions, we introduce a Distribution-based Acquisition approach that makes minimal assumption regarding the data pool and facilitates the data acquisition across various settings. Through extensive experimentation encompassing diverse datasets, models, and parameter configurations, we demonstrate the efficacy of our proposed methods across a range of tasks. Comparative experiments with alternative applicable baselines underscore the superior performance of our proposed approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a6edff53-9703-46d5-bedd-6b8c8f82a2f2Builds on15
- Mix-n-Match : Ensemble and Compositional Methods for Uncertainty Calibration in Deep LearningJize Zhang, Bhavya Kailkhura, Thomas Yong-Jin HanICML 2020 · 276 citations
- Rethinking Bias-Variance Trade-off for Generalization of Neural NetworksZitong Yang, Yaodong Yu, Chong You, Jacob Steinhardt et al.ICML 2020 · 219 citations
- DDG-DA: Data Distribution Generation for Predictable Concept Drift AdaptationWendi Li, Xiao Yang, Weiqing Liu, Yingce Xia et al.AAAI 2022 · 79 citations
- Energy-Based Open-World Uncertainty Modeling for Confidence CalibrationYezhen Wang, Bo Li, Tong Che, Kaiyang Zhou et al.ICCV 2021 · 78 citations
- Selective Data Acquisition in the Wild for Model ChargingChengliang Chai, Jiabin Liu, Nan Tang, Guoliang Li et al.VLDB 2022 · 62 citations
Related papers
- Data Acquisition for Improving Machine Learning ModelsYifan Li, Xiaohui Yu, Nick KoudasVLDB 2021 · 57 citations
- Optimizing Data Acquisition to Enhance Machine Learning PerformanceTingting Wang, Shixun Huang, Zhifeng Bao, J. Shane Culpepper et al.VLDB 2024 · 13 citations
- Data Acquisition via Experimental Design for Data MarketsCharles Lu, Baihe Huang, Sai Praneeth Karimireddy, Praneeth Vepakomma et al.NeurIPS 2024 · 12 citations
- Generator Assisted Mixture of Experts for Feature Acquisition in BatchVedang Asgaonkar, Aditya Jain, Abir DeAAAI 2024 · 3 citations
- Train on Validation (ToV): Fast data selection with applications to fine-tuningAyush Jain, Andrea Montanari, Eren SasogluICLR 2026 · 4 citations
