Retiring Adult: New Datasets for Fair Machine Learning
Frances Ding, Moritz Hardt, John Miller, Ludwig Schmidt
Abstract
Although the fairness community has recognized the importance of data, researchers in the area primarily rely on UCI Adult when it comes to tabular data. Derived from a 1994 US Census survey, this dataset has appeared in hundreds of research papers where it served as the basis for the development and comparison of many algorithmic fairness interventions. We reconstruct a superset of the UCI Adult data from available US Census sources and reveal idiosyncrasies of the UCI Adult dataset that limit its external validity. Our primary contribution is a suite of new datasets derived from US Census surveys that extend the existing data ecosystem for research on fair machine learning. We create prediction tasks relating to income, employment, health, transportation, and housing. The data span multiple years and all states of the United States, allowing researchers to study temporal shift and geographic variation. We highlight a broad initial sweep of new empirical insights relating to trade-offs between fairness criteria, performance of algorithmic interventions, and the role of distribution shift based on our new datasets. Our findings inform ongoing debates, challenge some existing narratives, and point to future research directions. Our datasets are available at https://github.com/zykls/folktables.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext acc13627-ff52-4c58-a3d9-780e18d24b9aCited by top-tier papers187
- Beyond Memorization: Violating Privacy via Inference with Large Language ModelsRobin Staab, Mark Vero, Mislav Balunovic, Martin T. VechevICLR 2024 · 211 citations
- Picking on the Same Person: Does Algorithmic Monoculture lead to Outcome Homogenization?Rishi Bommasani, Kathleen A. Creel, Ananya Kumar, Dan Jurafsky et al.NeurIPS 2022 · 179 citations
- Understanding the Role of Human Intuition on Reliance in Human-AI Decision-Making with ExplanationsValerie Chen, Q. Vera Liao, Jennifer Wortman Vaughan, Gagan BansalCSCW 2023 · 146 citations
- Questioning the Survey Responses of Large Language ModelsRicardo Dominguez-Olmedo, Moritz Hardt, Celestine Mendler-DünnerNeurIPS 2024 · 116 citations
- Scalable Membership Inference Attacks via Quantile RegressionMartin Bertran Lopez, Shuai Tang, Aaron Roth, Michael Kearns et al.NeurIPS 2023 · 96 citations
Related papers
- Benchmarking Stochastic Approximation Algorithms for Fairness-Constrained Training of Deep Neural NetworksAndrii Kliachkin, Jana Lepsová, Gilles Bareilles, Jakub MarecekICLR 2026 · 1 citation
- Social Bias Meets Data Bias: The Impacts of Labeling and Measurement Errors on Fairness CriteriaYiqiao Liao, Parinaz NaghizadehAAAI 2023 · 15 citations
- Fairness Guarantees under Demographic ShiftStephen Giguere, Blossom Metevier, Bruno Castro da Silva, Yuriy Brun et al.ICLR 2022 · 57 citations
- On the Fairness of Causal Algorithmic RecourseJulius von Kügelgen, Amir-Hossein Karimi, Umang Bhatt, Isabel Valera et al.AAAI 2022 · 99 citations
- Fairness Transferability Subject to Bounded Distribution ShiftYatong Chen, Reilly Raab, Jialu Wang, Yang LiuNeurIPS 2022 · 40 citations
