Auditing for Diversity Using Representative Examples
Vijay Keswani, L. Elisa Celis
摘要
Assessing the diversity of a dataset of information associated with people is crucial before using such data for downstream applications. For a given dataset, this often involves computing the imbalance or disparity in the empirical marginal distribution of a protected attribute (e.g. gender, dialect, etc.). However, real-world datasets, such as images from Google Search or collections of Twitter posts, often do not have protected attributes labeled. Consequently, to derive disparity measures for such datasets, the elements need to hand-labeled or crowd-annotated, which are expensive processes. We propose a cost-effective approach to approximate the disparity of a given unlabeled dataset, with respect to a protected attribute, using a control set of labeled representative examples. Our proposed algorithm uses the pairwise similarity between elements in the dataset and elements in the control set to effectively bootstrap an approximation to the disparity of the dataset. Importantly, we show that using a control set whose size is much smaller than the size of the dataset is sufficient to achieve a small approximation error. Further, based on our theoretical framework, we also provide an algorithm to construct adaptive control sets that achieve smaller approximation errors than randomly chosen control sets. Simulations on two image datasets and one Twitter dataset demonstrate the efficacy of our approach (using random and adaptive control sets) in auditing the diversity of a wide variety of datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper2
相关 Paper
- Estimation of Fair Ranking Metrics with Incomplete JudgmentsÖmer Kirnap, Fernando Diaz, Asia Biega, Michael D. Ekstrand 等WWW 2021 · 被引用 40 次
- Dialect Diversity in Text Summarization on TwitterVijay Keswani, L. Elisa CelisWWW 2021 · 被引用 26 次
- Estimating Structural Disparities for Face ModelsShervin Ardeshir, Cristina Segalin, Nathan KallusCVPR 2022 · 被引用 2 次
- Label-Efficient Group Robustness via Out-of-Distribution Concept CurationYiwei Yang, Anthony Z. Liu, Robert Wolfe, Aylin Caliskan 等CVPR 2024
- Controllable Guarantees for Fair Outcomes via Contrastive Information EstimationUmang Gupta, Aaron M. Ferber, Bistra Dilkina, Greg Ver SteegAAAI 2021 · 被引用 78 次
