NoiseGPT: Label Noise Detection and Rectification through Probability Curvature
Haoyu Wang, Zhuo Huang, Zhiwei Lin, Tongliang Liu
Abstract
Machine learning craves high-quality data which is a major bottleneck during realistic deployment, as it takes abundant resources and massive human labor to collect and label data. Unfortunately, label noise where image data mismatches with incorrect label exists ubiquitously in all kinds of datasets, significantly degrading the learning performance of deep networks. Learning with Label Noise (LNL) has been a common strategy for mitigating the influence of noisy labels. However, existing LNL methods either require pertaining using the memorization effect to separate clean data from noisy ones or rely on dataset assumptions that cannot extend to various scenarios. Thanks to the development of Multimodal Large Language Models (MLLMs) which possess massive knowledge and hold In-Context Learning (ICL) ability, this paper proposes NoiseGPT to effectively leverage MLLMs as a knowledge expert for conducting label noise detection and rectification. Specifically, we observe a probability curvature effect of MLLMs where clean and noisy examples reside on curvatures with different smoothness, further enabling the detection of label noise. By designing a token-wise Mix-of-Feature (MoF) technique to produce the curvature, we propose an In-Context Discrepancy (ICD) measure to determine the authenticity of an image-label pair. Subsequently, we repeat such a process to find the best matching pairs to complete our label rectification. Through extensive experiments, we carefully demonstrate the effectiveness of NoiseGPT on detecting and cleansing dataset noise, especially on ILSVRC12, the AUROC of NoiseGPT reached over 0.92. And by integrating with existing methods, the classification performance can be significantly improved on noisy datasets, typically by 22.8% on 80% symmetric CIFAR-10 with M-correction
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1bd89f86-fd85-486f-9cac-9fadeda97708Cited by top-tier papers11
- L-MTP: Leap Multi-Token Prediction Beyond Adjacent Context for Large Language ModelsXiaohao Liu, Xiaobo Xia, Weixiang Zhao, Manyi Zhang et al.NeurIPS 2025 · 16 citations
- FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEctionRunqi Lin, Alasdair Paren, Suqin Yuan, Muyang Li et al.CVPR 2026 · 13 citations
- Debiased Sample Selection for Learning with Noisy LabelsWeiran Pan, Wei Wei, Wenfeng XieCVPR 2026
- On Revisiting Entropy for Identifying Mislabeled ImagesChunlei Li, Zixuan Zheng, Yilei Shi, Guanglu Dong et al.ICML 2026
- TANGO: Text-Anchored Guided Optimization for Robust Fine-tuning Vision-Language Models under Label NoiseTengfei Ma, Weiran Pan, Wei WeiCVPR 2026
Builds on26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Meta Label Correction for Noisy Label LearningGuoqing Zheng, Ahmed Hassan Awadallah, Susan T. DumaisAAAI 2021 · 239 citations
- Hide and Seek in Noise Labels: Noise-Robust Collaborative Active Learning with LLMs-Powered AssistanceBo Yuan, Yulin Chen, Yin Zhang, Wei JiangACL 2024
- Mitigating Endogenous Confirmation Bias in Noisy Label Learning for Vision-Language ModelsFeiyang Ning, Xinyang ChenAAAI 2026
- Fuzzy Learning MachineJunbiao Cui, Jiye LiangNeurIPS 2022 · 6 citations
- Label Noise Correction via Fuzzy Learning MachineJiye Liang, Yixiao Li, Junbiao CuiAAAI 2025 · 3 citations
