Understanding Bias in Large-Scale Visual Datasets
Boya Zeng, Yida Yin, Zhuang Liu
Abstract
A recent study has shown that large-scale visual datasets are very biased: they can be easily classified by modern neural networks. However, the concrete forms of bias among these datasets remain unclear. In this study, we propose a framework to identify the unique visual attributes distinguishing these datasets. Our approach applies various transformations to extract semantic, structural, boundary, color, and frequency information from datasets, and assess how much each type of information reflects their bias. We further decompose their semantic bias with object-level analysis, and leverage natural language methods to generate detailed, open-ended descriptions of each dataset's characteristics. Our work aims to help researchers understand the bias in existing large-scale pre-training datasets, and build more diverse and representative ones in the future. Our project page and code are available at http://boyazeng.github.io/understand_bias .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Drag-and-Drop LLMs: Zero-Shot Prompt-to-WeightsZhiyuan Liang, Dongwen Tang, Yuhao Zhou, Xuanlei Zhao et al.NeurIPS 2025 · 22 citations
- ConceptScope: Characterizing Dataset Bias via Disentangled Visual ConceptsJinho Choi, Hyesu Lim, Steffen Schneider, Jaegul ChooNeurIPS 2025 · 5 citations
- PRISM: Reducing Spurious Implicit Biases in Vision-Language Models with LLM-Guided Embedding ProjectionMahdiyar Molahasani, Azadeh Motamedi, Michael A. Greenspan, Il-Min Kim et al.ICCV 2025 · 5 citations
- Revisiting Data Auditing in Large Vision-Language ModelsHongyu Zhu, Sichu Liang, Wenwen Wang, Boheng Li et al.ACM MM 2025 · 3 citations
- FedFACT: A Provable Framework for Controllable Group-Fairness Calibration in Federated LearningLi Zhang, Zhongxuan Han, Xiaohua Feng, Jiaming Zhang et al.NeurIPS 2025 · 2 citations
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
Related papers
- Classifier-to-Bias: Toward Unsupervised Automatic Bias Detection for Visual ClassifiersQuentin Guimard, Moreno D'Incà, Massimiliano Mancini, Elisa RicciCVPR 2025
- HiBug: On Human-Interpretable Model DebugMuxi Chen, Yu Li, Qiang XuNeurIPS 2023 · 22 citations
- BiasEdit: A Training-Free Bias-Detect-and-Edit Framework for Learning Fair Visual ClassifiersJungwook Seo, Yoonsik Park, Changmin Lee, Sungyong BaikWWW 2026
- A Decade's Battle on Dataset Bias: Are We There Yet?Zhuang Liu, Kaiming HeICLR 2025 · 8 citations
- Explaining Deep Convolutional Neural Networks via Latent Visual-Semantic Filter AttentionYu Yang, Seungbae Kim, Jungseock JooCVPR 2022 · 11 citations
