Mitigating Endogenous Confirmation Bias in Noisy Label Learning for Vision-Language Models
Feiyang Ning, Xinyang Chen
Abstract
Pretrained vision-language models (VLMs), especially CLIP, excel at adapting to downstream tasks through fine-tuning with sufficient high-quality labeled data. However, real-world training data often contains noisy labels, leading to significant performance degradation when models are naively fine-tuned on them. Existing noisy label learning methods for VLMs typically leverage the model's own pretrained knowledge, either via zero-shot predictions or vanilla self-training based on them, to identify and handle noisy samples. Crucially, these approaches blindly trust the VLM's pretrained knowledge, which can introduce endogenous confirmation bias: erroneous pretrained priors lead to incorrect noise detection, further amplifying the bias and corrupting the model. To overcome this limitation, we propose the Debiased Knowledge Adaptation Framework (DKAF), which empowers the model to challenge and correct potentially flawed zero-shot predictions. DKAF operates in three progressive phases: (1) Clean Sample Selection. We introduce a cross-modal collaborative pseudo-labeling to train a robust noisy label detector, explicitly mitigating confirmation bias by aggregating diverse signals beyond the model's initial zero-shot view. (2) Noisy Label Refinement. For samples identified as noisy, we apply a dual-modal consistency strategy to selectively correct their labels, leveraging alignment between dominant and fused modalities to guide refinement while minimizing reliance on potentially biased internal knowledge. (3) Model Adaptation. The model is progressively fine-tuned using the jointly curated dataset of selected clean samples and corrected noisy samples, promoting robust adaptation to the target task. Extensive experiments on nine benchmark datasets (both synthetic and real-world noise) demonstrate that DKAF consistently outperforms state-of-the-art multimodal noisy label learning methods. Notably, under high-noise conditions, DKAF achieves average accuracy improvements of 3.08%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo LabelingBowen Zhang, Yidong Wang, Wenxin Hou, Hao Wu et al.NeurIPS 2021 · 1,389 citations
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 1,326 citations
Related papers
- TrustCLIP: Learning from Noisy Labels via Semantic Label Verification and Trust-aligned Gradient ProjectionXueyi Zhang, Peiyin Zhu, Yuan Liao, Xiyu Wang et al.ACM MM 2025
- CLIPCleaner: Cleaning Noisy Labels with CLIPChen Feng, Georgios Tzimiropoulos, Ioannis PatrasACM MM 2024 · 12 citations
- Pre-Trained Vision-Language Models as Noisy Partial AnnotatorsQian-Wei Wang, Yuqiu Xie, Letian Zhang, Zimo Liu et al.AAAI 2025 · 3 citations
- Unbiased Region-Language Alignment for Open-Vocabulary Dense PredictionYunheng Li, Yuxuan Li, Quan-Sheng Zeng, Wenhai Wang et al.ICCV 2025 · 3 citations
- Vision-Language Models are Strong Noisy Label DetectorsTong Wei, Hao-Tian Li, Chun-Shu Li, Jiang-Xin Shi et al.NeurIPS 2024 · 26 citations
