FairDeDup: Detecting and Mitigating Vision-Language Fairness Disparities in Semantic Dataset Deduplication
Eric Slyman, Stefan Lee, Scott Cohen, Kushal Kafle
摘要
Recent dataset deduplication techniques have demonstrated that content-aware dataset pruning can dramatically reduce the cost of training Vision-Language Pre-trained (VLP) models without significant performance losses compared to training on the original dataset. These results have been based on pruning commonly used image-caption datasets collected from the web - datasets that are known to harbor harmful social biases that may then be codified in trained models. In this work, we evaluate how deduplication affects the prevalence of these biases in the resulting trained models and introduce an easy-to-implement modification to the recent SemDeDup algorithm that can reduce the negative effects that we observe. When examining CLIP-style models trained on deduplicated variants of LAION-400M, we find our proposed FairDeDup algorithm consistently leads to improved fairness metrics over SemDeDup on the FairFace and FACET datasets while maintaining zero-shot performance on CLIP benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- A Unified Debiasing Approach for Vision-Language Models across Modalities and TasksHoin Jung, Taeuk Jang, Xiaoqian WangNeurIPS 2024 · 被引用 27 次
- Scaling-Aware Data Selection for End-to-End Autonomous Driving SystemsTolga Dimlioglu, Nadine Chang, Maying Shen, Rafid Mahmood 等CVPR 2026 · 被引用 1 次
- Bias in Gender Bias Benchmarks: How Spurious Features Distort EvaluationYusuke Hirota, Ryo Hachiuma, Boyi Li, Ximing Lu 等ICCV 2025
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
相关 Paper
- Joint Vision-Language Social Bias Removal for CLIPHaoyu Zhang, Yangyang Guo, Mohan S. KankanhalliCVPR 2025
- Counterfactually Measuring and Eliminating Social Bias in Vision-Language Pre-training ModelsYi Zhang, Junyang Wang, Jitao SangACM MM 2022 · 被引用 11 次
- Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to ModelsLeander Girrbach, Stephan Alaniz, Genevieve Smith, Trevor Darrell 等ICLR 2026 · 被引用 5 次
- FairerCLIP: Debiasing CLIP's Zero-Shot Predictions using Functions in RKHSsSepehr Dehdashtian, Lan Wang, Vishnu BoddetiICLR 2024 · 被引用 34 次
- FairCLIP: Harnessing Fairness in Vision-Language LearningYan Luo, Min Shi, Muhammad Osama Khan, Muhammad Muneeb Afzal 等CVPR 2024 · 被引用 37 次
