Visual Data Diagnosis and Debiasing with Concept Graphs
Rwiddhi Chakraborty, Yinong Wang, Jialu Gao, Runkai Zheng, Cheng Zhang, Fernando De la Torre
Abstract
The widespread success of deep learning models today is owed to the curation of extensive datasets significant in size and complexity. However, such models frequently pick up inherent biases in the data during the training process, leading to unreliable predictions. Diagnosing and debiasing datasets is thus a necessity to ensure reliable model performance. In this paper, we present ConBias, a novel framework for diagnosing and mitigating Concept co-occurrence Biases in visual datasets. ConBias represents visual datasets as knowledge graphs of concepts, enabling meticulous analysis of spurious concept co-occurrences to uncover concept imbalances across the whole dataset. Moreover, we show that by employing a novel clique-based concept balancing strategy, we can mitigate these imbalances, leading to enhanced performance on downstream tasks. Extensive experiments show that data augmentation based on a balanced concept distribution augmented by Conbias improves generalization performance across multiple datasets compared to state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b23b1e74-d9a7-4988-addd-e4ce19f064c1Cited by top-tier papers2
- ConceptScope: Characterizing Dataset Bias via Disentangled Visual ConceptsJinho Choi, Hyesu Lim, Steffen Schneider, Jaegul ChooNeurIPS 2025 · 5 citations
- Representation-Level Counterfactual Calibration for Debiased Zero-Shot RecognitionPei Peng, Ming-Kun Xie, Hang Hao, Tong Jin et al.NeurIPS 2025 · 2 citations
Builds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- Common Sense Bias Modeling for Classification TasksMiao Zhang, Zee Fryer, Ben Colman, Ali Shahriyari et al.AAAI 2025
- CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic TriplesKyohoon Jin, Juhwan Choi, Jungmin Yun, Junho Lee et al.EMNLP 2025
- Counterexample Contrastive Learning for Spurious Correlation EliminationJinqiang Wang, Rui Hu, Chaoquan Jiang, Rui Hu et al.ACM MM 2022 · 3 citations
- Unbiased Scene Graph Generation From Biased TrainingKaihua Tang, Yulei Niu, Jianqiang Huang, Jiaxin Shi et al.CVPR 2020
- WRING Out The Bias: A Rotation-Based Alternative To Projection DebiasingWalter Gerych, Cassandra Parent, Quinn Perian, Rafiya Javed et al.ICLR 2026
