Towards Transparency in Dermatology Image Datasets with Skin Tone Annotations by Experts, Crowds, and an Algorithm
Matthew Groh, Caleb Harris, Roxana Daneshjou, Omar Badri, Arash Koochek
摘要
While artificial intelligence (AI) holds promise for supporting healthcare providers and improving the accuracy of medical diagnoses, a lack of transparency in the composition of datasets exposes AI models to the possibility of unintentional and avoidable mistakes.
In particular, public and private image datasets of dermatological conditions rarely include information on skin color. As a start towards increasing transparency, AI researchers have appropriated the use of the Fitzpatrick skin type (FST) from a measure of patient photosensitivity to a measure for estimating skin tone in algorithmic audits of computer vision applications including facial recognition and dermatology diagnosis. In order to understand the variability of estimated FST annotations on images, we compare several FST annotation methods on a diverse set of 460 images of skin conditions from both textbooks and online dermatology atlases.
These methods include expert annotation by board-certified dermatologists, algorithmic annotation via the Individual Typology Angle algorithm, which is then converted to estimated FST (ITA-FST), and two crowd-sourced, dynamic consensus protocols for annotating estimated FSTs. We find the inter-rater reliability between three board-certified dermatologists is comparable to the inter-rater reliability between the board-certified dermatologists and either of the crowdsourcing methods. In contrast, we find that the ITA-FST method produces annotations that are significantly less correlated with the experts' annotations than the experts' annotations are correlated with each other. These results demonstrate that algorithms based on ITA-FST are not reliable for annotating large-scale image datasets, but human-centered, crowd-based protocols can reliably add skin type transparency to dermatology datasets. Furthermore, we introduce the concept of dynamic consensus protocols with tunable parameters including expert review that increase the visibility of crowdwork and provide guidance for future crowdsourced annotations of large image datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- FACET: Fairness in Computer Vision Evaluation BenchmarkLaura Gustafson, Chloé Rolland, Nikhila Ravi, Quentin Duval 等ICCV 2023 · 被引用 74 次
- FairTune: Optimizing Parameter Efficient Fine Tuning for Fairness in Medical Image AnalysisRaman Dutt, Ondrej Bohdal, Sotirios A. Tsaftaris, Timothy M. HospedalesICLR 2024 · 被引用 30 次
- Beyond Skin Tone: A Multidimensional Measure of Apparent Skin ColorWilliam Thong, Przemyslaw Joniak, Alice XiangICCV 2023 · 被引用 29 次
- Technical Responses To Critique: The Case Of Skin ToneSayan Bhattacharjee, David RibesCHI 2025 · 被引用 2 次
- Learning to Be a Doctor: Searching for Effective Medical Agent ArchitecturesYangyang Zhuang, Wenjia Jiang, Jiayu Zhang, Ze Yang 等ACM MM 2025 · 被引用 1 次
它引用的顶会 Paper14
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong 等CHI 2021 · 被引用 725 次
- A Human-Centered Evaluation of a Deep Learning System Deployed in Clinics for the Detection of Diabetic RetinopathyEmma Beede, Elizabeth Elliott Baylor, Fred Hersch, Anna Iurchenko 等CHI 2020 · 被引用 589 次
- Expanding Explainability: Towards Social Transparency in AI systemsUpol Ehsan, Q. Vera Liao, Michael J. Muller, Mark O. Riedl 等CHI 2021 · 被引用 505 次
- How to Evaluate Trust in AI-Assisted Decision Making? A Survey of Empirical MethodologiesOleksandra Vereschak, Gilles Bailly, Baptiste CaramiauxCSCW 2021 · 被引用 227 次
- How We've Taught Algorithms to See Identity: Constructing Race and Gender in Image Databases for Facial AnalysisMorgan Klaus Scheuerman, Kandrea Wade, Caitlin Lustig, Jed R. BrubakerCSCW 2020 · 被引用 198 次
相关 Paper
- Toward a Perspectivist Turn in Ground Truthing for Predictive ComputingFederico Cabitza, Andrea Campagner, Valerio BasileAAAI 2023 · 被引用 236 次
- Benchmarking Algorithmic Bias in Face Recognition: An Experimental Approach Using Synthetic Faces and Human EvaluationHao Liang, Pietro Perona, Guha BalakrishnanICCV 2023 · 被引用 33 次
- It's About Time: A View of Crowdsourced Data Before and During the PandemicEvgenia Christoforou, Pinar Barlas, Jahna OtterbacherCHI 2021 · 被引用 13 次
- MICA: Towards Explainable Skin Lesion Diagnosis via Multi-Level Image-Concept AlignmentYequan Bie, Luyang Luo, Hao ChenAAAI 2024 · 被引用 28 次
- Is this AI trained on Credible Data? The Effects of Labeling Quality and Performance Bias on User TrustCheng Chen, S. Shyam SundarCHI 2023 · 被引用 36 次
