FACET: Fairness in Computer Vision Evaluation Benchmark
Laura Gustafson, Chloé Rolland, Nikhila Ravi, Quentin Duval, Aaron Adcock, Cheng-Yang Fu, Melissa Hall, Candace Ross
摘要
Computer vision models have known performance disparities across attributes such as gender and skin tone. This means during tasks such as classification and detection, model performance differs for certain classes based on the demographics of the people in the image. These disparities have been shown to exist, but until now there has not been a unified approach to measure these differences for common use-cases of computer vision models. We present a new benchmark named FACET (FAirness in Computer Vision EvaluaTion), a large, publicly available evaluation set of 32k images for some of the most common vision tasks -image classification, object detection and segmentation. For every image in FACET, we hired expert reviewers to manually annotate person-related attributes such as perceived skin tone and hair type, manually draw bounding boxes and label fine-grained person-related classes such as disk jockey or guitarist. In addition, we use FACET to benchmark state-of-the-art vision models and present a deeper understanding of potential performance disparities and challenges across sensitive demographic attributes. With the exhaustive annotations collected, we probe models using single demographics attributes as well as multiple attributes using an intersectional approach (e.g. hair color and perceived skin tone). Our results show that classification, detection, segmentation, and visual grounding models exhibit performance disparities across demographic attributes and intersections of attributes. These harms suggest that not all people represented in datasets receive fair and equitable treatment in these vision tasks. We hope current and future results using our benchmark will contribute to fairer, more robust vision models. FACET is available publicly at https://facet.metademolab.com . Dataset Dataset Size Apparent or Self-Reported Attributes Task #/people #/images #/videos #/boxes #/masks gender age skin tone race lighting additional UTK-Face[98] 20k 20k ---Yes Yes No Yes No No -FairFace[57] 108k 108k ---Yes Yes No Yes No No -Gender Shades[8] 1.2k 1.2k ---Yes Yes Yes No No No -OpenImages MIAP[84] 454k 100k -454k * Yes Yes No No No No C * DS * [94] annotations for BDDK 100k [97] 16k 2.2k -16k * No No Yes No Yes No DS * [100] annotations for COCO [63] 28k 16k -28k 28k Yes No Yes No No No C * DS Casual Conversations v1[43] 3k N/A 45k --Yes Yes Yes No Yes Yes -Casual Conversations v2 [42] 5.6k N/A 26k --Yes Yes Yes No Yes Yes -Ours -FACET 50k 32k -50k 69k Yes Yes Yes No Yes Yes CDS * represents tasks/annotations that are not included in the fairness portion of the dataset, but are included in the overall dataset. e.g COCO has been used for multi-class classification [101, 92]
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper28
- A Unified Debiasing Approach for Vision-Language Models across Modalities and TasksHoin Jung, Taeuk Jang, Xiaoqian WangNeurIPS 2024 · 被引用 27 次
- Image Clustering Conditioned on Text CriteriaSehyun Kwon, Jaeseung Park, Minkyu Kim, Jaewoong Cho 等ICLR 2024 · 被引用 27 次
- ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image GenerationAkshita Jha, Vinodkumar Prabhakaran, Remi Denton, Sarah Laszlo 等ACL 2024 · 被引用 18 次
- Visual Data Diagnosis and Debiasing with Concept GraphsRwiddhi Chakraborty, Yinong Wang, Jialu Gao, Runkai Zheng 等NeurIPS 2024 · 被引用 9 次
- Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post‑hoc Debiasing in Vision-Language ModelsDachuan Zhao, Weiyue Li, Zhenda Shen, Yushu Qiu 等CVPR 2026 · 被引用 5 次
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong 等CHI 2021 · 被引用 725 次
- How We've Taught Algorithms to See Identity: Constructing Race and Gender in Image Databases for Facial AnalysisMorgan Klaus Scheuerman, Kandrea Wade, Caitlin Lustig, Jed R. BrubakerCSCW 2020 · 被引用 198 次
- Understanding and Evaluating Racial Biases in Image CaptioningDora Zhao, Angelina Wang, Olga RussakovskyICCV 2021 · 被引用 165 次
相关 Paper
- Benchmarking Algorithmic Bias in Face Recognition: An Experimental Approach Using Synthetic Faces and Human EvaluationHao Liang, Pietro Perona, Guha BalakrishnanICCV 2023 · 被引用 33 次
- Leveraging Diffusion Perturbations for Measuring Fairness in Computer VisionNicholas Lui, Bryan Chia, William Berrios, Candace Ross 等AAAI 2024 · 被引用 3 次
- Bias in Gender Bias Benchmarks: How Spurious Features Distort EvaluationYusuke Hirota, Ryo Hachiuma, Boyi Li, Ximing Lu 等ICCV 2025
- Measuring Social Bias in Vision-Language Models with Face-Only Counterfactuals from Real PhotosHaodong Chen, Qiang Huang, Jiaqi Zhao, Qiuping Jiang 等ACL 2026 · 被引用 1 次
- Beyond Skin Tone: A Multidimensional Measure of Apparent Skin ColorWilliam Thong, Przemyslaw Joniak, Alice XiangICCV 2023 · 被引用 29 次
