Falcon: Fair Active Learning using Multi-armed Bandits
Ki Hyun Tae, Hantian Zhang, Jaeyoung Park, Kexin Rong, Steven Euijong Whang
Abstract
Biased data can lead to unfair machine learning models, highlighting the importance of embedding fairness at the beginning of data analysis, particularly during dataset curation and labeling. In response, we propose Falcon, a scalable fair active learning framework. Falcon adopts a data-centric approach that improves machine learning model fairness via strategic sample selection. Given a user-specified group fairness measure, Falcon identifies samples from "target groups" (e.g., (attribute=female, label=positive)) that are the most informative for improving fairness. However, a challenge arises since these target groups are defined using ground truth labels that are not available during sample selection. To handle this, we propose a novel trial-and-error method, where we postpone using a sample if the predicted label is different from the expected one and falls outside the target group. We also observe the trade-off that selecting more informative samples results in higher likelihood of postponing due to undesired label prediction, and the optimal balance varies per dataset. We capture the trade-off between informativeness and postpone rate as policies and propose to automatically select the best policy using adversarial multi-armed bandit methods, given their computational efficiency and theoretical guarantees. Experiments show that Falcon significantly outperforms existing fair active learning approaches in terms of fairness and accuracy and is more efficient. In particular, only Falcon supports a proper trade-off between accuracy and fairness where its maximum fairness score is 1.8--4.5x higher than the second-best results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f09cca43-13a6-4f34-9a06-3a5512d27409Cited by top-tier papers6
- DRMD: Deep Reinforcement Learning for Malware Detection Under Concept DriftShae McFadden, Myles Foley, Mario D'Onghia, Chris Hicks et al.AAAI 2026 · 7 citations
- Low Rank Learning for Offline Query OptimizationZixuan Yi, Yao Tian, Zachary G. Ives, Ryan MarcusSIGMOD 2025 · 4 citations
- Fair and Actionable Causal Prescription RulesetBenton Li, Nativ Levy, Brit Youngmann, Sainyam Galhotra et al.SIGMOD 2025 · 3 citations
- Fair Data Pre-Processing with Imperfect Attribute SpaceYing Zheng, Yangfan Jiang, Kian-Lee TanSIGMOD 2026
- CausalPre: Scalable and Effective Data Pre-Processing for Causal FairnessYing Zheng, Yangfan Jiang, Kian-Lee TanICDE 2026
Builds on16
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford et al.ICLR 2020 · 974 citations
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 671 citations
- Is There a Trade-Off Between Fairness and Accuracy? A Perspective Using Mismatched Hypothesis TestingSanghamitra Dutta, Dennis Wei, Hazar Yueksel, Pin-Yu Chen et al.ICML 2020 · 171 citations
- FairBatch: Batch Selection for Model FairnessYuji Roh, Kangwook Lee, Steven Euijong Whang, Changho SuhICLR 2021 · 156 citations
- Operationalizing Individual Fairness with Pairwise Fair RepresentationsPreethi Lahoti, Krishna P. Gummadi, Gerhard WeikumVLDB 2020 · 88 citations
Related papers
- Fairness without Harm: An Influence-Guided Active Sampling ApproachJinlong Pang, Jialu Wang, Zhaowei Zhu, Yuanshun Yao et al.NeurIPS 2024 · 13 citations
- Fairness-Aware Active Online Learning with Changing EnvironmentsSadaf Md. Halim, Chen Zhao, Xintao Wu, Latifur Khan et al.ICDE 2025 · 1 citation
- Fair Bayesian Data Selection via Generalized Discrepancy MeasuresYixuan Zhang, Jiabin Luo, Zhenggang Wang, Feng Zhou et al.AAAI 2026
- Active Statistical InferenceTijana Zrnic, Emmanuel J. CandèsICML 2024 · 34 citations
- SEL-BALD: Deep Bayesian Active Learning with Selective LabelsRuijiang Gao, Mingzhang Yin, Maytal Saar-TsechanskyNeurIPS 2024 · 4 citations
