Data Acquisition for Improving Model Confidence
Yifan Li, Xiaohui Yu, Nick Koudas
摘要
In recent years, there has been a growing recognition that high-quality training data is crucial for the performance of machine learning models. This awareness has catalyzed both research endeavors and industrial initiatives dedicated to data acquisition to enhance diverse dimensions of model performance. Among these dimensions, model confidence holds paramount importance; however, it has often been overlooked in prior investigations into data acquisition methodologies. To address this gap, our work focuses on improving the data acquisition process with the goal of enhancing the confidence of Machine Learning models. Specifically, we operate within a practical context where limited samples can be obtained from a large data pool. We employ well-established model confidence metrics as our foundation, and we propose two methodologies, Bulk Acquisition (BA) and Sequential Acquisition (SA), each geared towards identifying the sets of samples that yield the most substantial gains in model confidence. Recognizing the complexity of BA and SA, we introduce two efficient approximate methods, namely kNN-BA and kNN-SA, restricting data acquisition to promising subsets within the data pool. To broaden the applicability of our solutions, we introduce a Distribution-based Acquisition approach that makes minimal assumption regarding the data pool and facilitates the data acquisition across various settings. Through extensive experimentation encompassing diverse datasets, models, and parameter configurations, we demonstrate the efficacy of our proposed methods across a range of tasks. Comparative experiments with alternative applicable baselines underscore the superior performance of our proposed approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Mix-n-Match : Ensemble and Compositional Methods for Uncertainty Calibration in Deep LearningJize Zhang, Bhavya Kailkhura, Thomas Yong-Jin HanICML 2020 · 被引用 276 次
- Rethinking Bias-Variance Trade-off for Generalization of Neural NetworksZitong Yang, Yaodong Yu, Chong You, Jacob Steinhardt 等ICML 2020 · 被引用 219 次
- DDG-DA: Data Distribution Generation for Predictable Concept Drift AdaptationWendi Li, Xiao Yang, Weiqing Liu, Yingce Xia 等AAAI 2022 · 被引用 79 次
- Energy-Based Open-World Uncertainty Modeling for Confidence CalibrationYezhen Wang, Bo Li, Tong Che, Kaiyang Zhou 等ICCV 2021 · 被引用 78 次
- Selective Data Acquisition in the Wild for Model ChargingChengliang Chai, Jiabin Liu, Nan Tang, Guoliang Li 等VLDB 2022 · 被引用 62 次
相关 Paper
- Data Acquisition for Improving Machine Learning ModelsYifan Li, Xiaohui Yu, Nick KoudasVLDB 2021 · 被引用 57 次
- Optimizing Data Acquisition to Enhance Machine Learning PerformanceTingting Wang, Shixun Huang, Zhifeng Bao, J. Shane Culpepper 等VLDB 2024 · 被引用 13 次
- Data Acquisition via Experimental Design for Data MarketsCharles Lu, Baihe Huang, Sai Praneeth Karimireddy, Praneeth Vepakomma 等NeurIPS 2024 · 被引用 12 次
- Generator Assisted Mixture of Experts for Feature Acquisition in BatchVedang Asgaonkar, Aditya Jain, Abir DeAAAI 2024 · 被引用 3 次
- Train on Validation (ToV): Fast data selection with applications to fine-tuningAyush Jain, Andrea Montanari, Eren SasogluICLR 2026 · 被引用 4 次
