Fair Classification with Partial Feedback: An Exploration-Based Data Collection Approach
Vijay Keswani, Anay Mehrotra, L. Elisa Celis
Abstract
In many predictive contexts (e.g., credit lending), true outcomes are only observed for samples that were positively classified in the past. These past observations, in turn, form training datasets for classifiers that make future predictions. However, such training datasets lack information about the outcomes of samples that were (incorrectly) negatively classified in the past and can lead to erroneous classifiers. We present an approach that trains a classifier using available data and comes with a family of exploration strategies to collect outcome data about subpopulations that otherwise would have been ignored. For any exploration strategy, the approach comes with guarantees that (1) all sub-populations are explored, (2) the fraction of false positives is bounded, and (3) the trained classifier converges to a ``desired'' classifier. The right exploration strategy is context-dependent; it can be chosen to improve learning guarantees and encode context-specific group fairness properties. Evaluation on real-world datasets shows that this approach consistently boosts the quality of collected outcome data and improves the fraction of true positives for all groups, with only a small reduction in predictive utility.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on12
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong et al.CHI 2021 · 725 citations
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 671 citations
- Robust Optimization for Fairness with Noisy Protected GroupsSerena Lutong Wang, Wenshuo Guo, Harikrishna Narasimhan, Andrew Cotter et al.NeurIPS 2020 · 134 citations
- Achieving Fairness in the Stochastic Multi-Armed Bandit ProblemVishakha Patil, Ganesh Ghalme, Vineet Nair, Y. NarahariAAAI 2020 · 131 citations
- Characterizing Fairness Over the Set of Good Models Under Selective LabelsAmanda Coston, Ashesh Rambachan, Alexandra ChouldechovaICML 2021 · 98 citations
Related papers
- Representation Matters: Assessing the Importance of Subgroup Allocations in Training DataEsther Rolf, Theodora T. Worledge, Benjamin Recht, Michael I. JordanICML 2021 · 50 citations
- Adaptive Data Debiasing through Bounded ExplorationYifan Yang, Yang Liu, Parinaz NaghizadehNeurIPS 2022 · 9 citations
- Fair Conformal Classification via Learning Representation-Based GroupsSenrong Xu, Yanke Zhou, Yuhao Tan, Zenan Li et al.ICLR 2026 · 1 citation
- The Importance of Modeling Data Missingness in Algorithmic Fairness: A Causal PerspectiveNaman Goel, Alfonso Amayuelas, Amit Deshpande, Amit SharmaAAAI 2021 · 36 citations
- Learning Models for Actionable RecourseAlexis Ross, Himabindu Lakkaraju, Osbert BastaniNeurIPS 2021 · 25 citations
