Retiring Adult: New Datasets for Fair Machine Learning
Frances Ding, Moritz Hardt, John Miller, Ludwig Schmidt
摘要
Although the fairness community has recognized the importance of data, researchers in the area primarily rely on UCI Adult when it comes to tabular data. Derived from a 1994 US Census survey, this dataset has appeared in hundreds of research papers where it served as the basis for the development and comparison of many algorithmic fairness interventions. We reconstruct a superset of the UCI Adult data from available US Census sources and reveal idiosyncrasies of the UCI Adult dataset that limit its external validity. Our primary contribution is a suite of new datasets derived from US Census surveys that extend the existing data ecosystem for research on fair machine learning. We create prediction tasks relating to income, employment, health, transportation, and housing. The data span multiple years and all states of the United States, allowing researchers to study temporal shift and geographic variation. We highlight a broad initial sweep of new empirical insights relating to trade-offs between fairness criteria, performance of algorithmic interventions, and the role of distribution shift based on our new datasets. Our findings inform ongoing debates, challenge some existing narratives, and point to future research directions. Our datasets are available at https://github.com/zykls/folktables.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper187
- Beyond Memorization: Violating Privacy via Inference with Large Language ModelsRobin Staab, Mark Vero, Mislav Balunovic, Martin T. VechevICLR 2024 · 被引用 211 次
- Picking on the Same Person: Does Algorithmic Monoculture lead to Outcome Homogenization?Rishi Bommasani, Kathleen A. Creel, Ananya Kumar, Dan Jurafsky 等NeurIPS 2022 · 被引用 179 次
- Understanding the Role of Human Intuition on Reliance in Human-AI Decision-Making with ExplanationsValerie Chen, Q. Vera Liao, Jennifer Wortman Vaughan, Gagan BansalCSCW 2023 · 被引用 146 次
- Questioning the Survey Responses of Large Language ModelsRicardo Dominguez-Olmedo, Moritz Hardt, Celestine Mendler-DünnerNeurIPS 2024 · 被引用 116 次
- Scalable Membership Inference Attacks via Quantile RegressionMartin Bertran Lopez, Shuai Tang, Aaron Roth, Michael Kearns 等NeurIPS 2023 · 被引用 96 次
相关 Paper
- Benchmarking Stochastic Approximation Algorithms for Fairness-Constrained Training of Deep Neural NetworksAndrii Kliachkin, Jana Lepsová, Gilles Bareilles, Jakub MarecekICLR 2026 · 被引用 1 次
- Social Bias Meets Data Bias: The Impacts of Labeling and Measurement Errors on Fairness CriteriaYiqiao Liao, Parinaz NaghizadehAAAI 2023 · 被引用 15 次
- Fairness Guarantees under Demographic ShiftStephen Giguere, Blossom Metevier, Bruno Castro da Silva, Yuriy Brun 等ICLR 2022 · 被引用 57 次
- On the Fairness of Causal Algorithmic RecourseJulius von Kügelgen, Amir-Hossein Karimi, Umang Bhatt, Isabel Valera 等AAAI 2022 · 被引用 99 次
- Fairness Transferability Subject to Bounded Distribution ShiftYatong Chen, Reilly Raab, Jialu Wang, Yang LiuNeurIPS 2022 · 被引用 40 次
