Vision-Language Models are Strong Noisy Label Detectors
Tong Wei, Hao-Tian Li, Chun-Shu Li, Jiang-Xin Shi, Yufeng Li, Min-Ling Zhang
摘要
Recent research on fine-tuning vision-language models has demonstrated impressive performance in various downstream tasks. However, the challenge of obtaining accurately labeled data in real-world applications poses a significant obstacle during the fine-tuning process. To address this challenge, this paper presents a Denoising Fine-Tuning framework, called DeFT, for adapting vision-language models. DeFT utilizes the robust alignment of textual and visual features pre-trained on millions of auxiliary image-text pairs to sieve out noisy labels. The proposed framework establishes a noisy label detector by learning positive and negative textual prompts for each class. The positive prompt seeks to reveal distinctive features of the class, while the negative prompt serves as a learnable threshold for separating clean and noisy samples. We employ parameter-efficient fine-tuning for the adaptation of a pre-trained visual encoder to promote its alignment with the learned textual prompts. As a general framework, DeFT can seamlessly fine-tune many pre-trained models to downstream tasks by utilizing carefully selected clean samples. Experimental results on seven synthetic and real-world noisy datasets validate the effectiveness of DeFT in both noisy label detection and image classification.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Handling Label Noise via Instance-Level Difficulty Modeling and Dynamic OptimizationKuan Zhang, Chengliang Chai, Jingzhe Xu, Chi Zhang 等NeurIPS 2025 · 被引用 6 次
- Nonparametric Teaching of Attention LearnersChen Zhang, Jianghui Wang, Bingyang Cheng, Zhongtao Chen 等ICLR 2026 · 被引用 3 次
- Debiased Sample Selection for Learning with Noisy LabelsWeiran Pan, Wei Wei, Wenfeng XieCVPR 2026
- On Revisiting Entropy for Identifying Mislabeled ImagesChunlei Li, Zixuan Zheng, Yilei Shi, Guanglu Dong 等ICML 2026
- TANGO: Text-Anchored Guided Optimization for Robust Fine-tuning Vision-Language Models under Label NoiseTengfei Ma, Weiran Pan, Wei WeiCVPR 2026
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
相关 Paper
- DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image ModelsKomal Kumar, Rao Muhammad Anwer, Fahad Shahbaz Khan, Salman H. Khan 等NeurIPS 2025 · 被引用 2 次
- Mitigating Endogenous Confirmation Bias in Noisy Label Learning for Vision-Language ModelsFeiyang Ning, Xinyang ChenAAAI 2026
- Exploring Cross-Modal Flows for Few-Shot LearningZiqi Jiang, Yanghao Wang, Long ChenICLR 2026 · 被引用 6 次
- AutoVP: An Automated Visual Prompting Framework and BenchmarkHsi-Ai Tsao, Lei Hsiung, Pin-Yu Chen, Si Liu 等ICLR 2024 · 被引用 29 次
- Fine-Grained Visual Prompt Learning of Vision-Language Models for Image RecognitionHongbo Sun, Xiangteng He, Jiahuan Zhou, Yuxin PengACM MM 2023 · 被引用 16 次
