Slice Tuner: A Selective Data Acquisition Framework for Accurate and Fair Machine Learning Models
Ki Hyun Tae, Steven Euijong Whang
Abstract
As machine learning becomes democratized in the era of Software 2.0, a serious bottleneck is acquiring enough data to ensure accurate and fair models. Recent techniques including crowdsourcing provide cost-effective ways to gather such data. However, simply acquiring data as much as possible is not necessarily an effective strategy for optimizing accuracy and fairness. For example, if an online app store has enough training data for certain slices of data (say American customers), but not for others, obtaining more American customer data will only bias the model training. Instead, we contend that one needs to selectively acquire data and propose Slice Tuner, which acquires possibly-different amounts of data per slice such that the model accuracy and fairness on all slices are optimized. This problem is different than labeling existing data (as in active learning or weak supervision) because the goal is obtaining the right amounts of new data. At its core, Slice Tuner maintains learning curves of slices that estimate the model accuracies given more data and uses convex optimization to find the best data acquisition strategy. The key challenges of estimating learning curves are that they may be inaccurate if there is not enough data, and there may be dependencies among slices where acquiring data for one slice influences the learning curves of others. We solve these issues by iteratively and efficiently updating the learning curves as more data is acquired. We evaluate Slice Tuner on real datasets using crowdsourcing for data acquisition and show that Slice Tuner significantly outperforms baselines in terms of model accuracy and fairness, even when the learning curves cannot be reliably estimated.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d93f2ae2-08d7-4bfb-855e-db7658f6f8cbCited by top-tier papers12
- FairBatch: Batch Selection for Model FairnessYuji Roh, Kangwook Lee, Steven Euijong Whang, Changho SuhICLR 2021 · 156 citations
- Mandoline: Model Evaluation under Distribution ShiftMayee F. Chen, Karan Goel, Nimit Sharad Sohoni, Fait Poms et al.ICML 2021 · 84 citations
- SAGA: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning ApplicationsShafaq Siddiqi, Roman Kern, Matthias BoehmSIGMOD 2024 · 24 citations
- Fairness-Aware Range Queries for Selecting Unbiased DataSuraj Shetiya, Ian P. Swift, Abolfazl Asudeh, Gautam DasICDE 2022 · 19 citations
- Fairness without Harm: An Influence-Guided Active Sampling ApproachJinlong Pang, Jialu Wang, Zhaowei Zhu, Yuanshun Yao et al.NeurIPS 2024 · 13 citations
Related papers
- Constructing a Fair Classifier with Generated Fair DataTaeuk Jang, Feng Zheng, Xiaoqian WangAAAI 2021 · 44 citations
- Falcon: Fair Active Learning using Multi-armed BanditsKi Hyun Tae, Hantian Zhang, Jaeyoung Park, Kexin Rong et al.VLDB 2024 · 7 citations
- SliceTeller: A Data Slice-Driven Approach for Machine Learning Model ValidationXiaoyu Zhang, Jorge Piazentin Ono, Huan Song, Liang Gou et al.IEEE VIS 2022 · 43 citations
- Fairness-Aware Active Online Learning with Changing EnvironmentsSadaf Md. Halim, Chen Zhao, Xintao Wu, Latifur Khan et al.ICDE 2025 · 1 citation
- Fair regression via plug-in estimator and recalibration with statistical guaranteesEvgenii Chzhen, Christophe Denis, Mohamed Hebiri, Luca Oneto et al.NeurIPS 2020 · 52 citations
