Noise-aware Learning from Web-crawled Image-Text Data for Image Captioning
Wooyoung Kang, Jonghwan Mun, Sungjun Lee, Byungseok Roh
Abstract
Image captioning is one of the straightforward tasks that can take advantage of large-scale web-crawled data which provides rich knowledge about the visual world for a captioning model. However, since web-crawled data contains image-text pairs that are aligned at different levels, the inherent noises (e.g., misaligned pairs) make it difficult to learn a precise captioning model. While the filtering strategy can effectively remove noisy data, it leads to a decrease in learnable knowledge and sometimes brings about a new problem of data deficiency. To take the best of both worlds, we propose a Noise-aware Captioning (NoC) framework, which learns rich knowledge from the whole web-crawled data while being less affected by the noises. This is achieved by the proposed alignment-level-controllable captioner, which is learned using alignment levels of the image-text pairs as a control signal during training. The alignment-level-conditioned training allows the model to generate high-quality captions by simply setting the control signal to the desired alignment level at inference time. An in-depth analysis shows the effectiveness of our framework in handling noise. With two tasks of zero-shot captioning and text-to-image retrieval using generated captions (i.e., self-retrieval), we also demonstrate our model can produce high-quality captions in terms of descriptiveness and distinctiveness. The code is available at https://github.com/kakaobrain/noc.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 55767ee6-8a10-4246-81e2-d22c4616f547Cited by top-tier papers16
- Noise-Aware Image Captioning with Progressively Exploring Mismatched WordsZhongtian Fu, Kefei Song, Luping Zhou, Yang YangAAAI 2024 · 36 citations
- Image Captioning with Multi-Context Synthetic DataFeipeng Ma, Yizhou Zhou, Fengyun Rao, Yueyi Zhang et al.AAAI 2024 · 22 citations
- Mixture-of-Scores: Robust Image-Text Data Valuation via Three Lines of CodeSitong Wu, Haoru Tan, Yukang Chen, Shaofeng Zhang et al.ICCV 2025 · 4 citations
- Noisy Correspondence Rectification via Asymmetric Similarity LearningYunbo Wang, YuJie Wu, Zhien Dai, Can Tian et al.AAAI 2025 · 4 citations
- Unlearning the Noisy Correspondence Makes CLIP More RobustHaochen Han, Alex Jinpeng Wang, Peijun Ye, Fangming LiuICCV 2025 · 3 citations
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- Noise-Aware Decoding with Salient Region Enhancing for Zero-Shot Image CaptioningYuxin Xie, Dongyue Chen, Yue Zhu, Tong Jia et al.ACM MM 2025
- Negative Entity Suppression for Zero-Shot Captioning with Synthetic ImagesZimao Lu, Hui Xu, Bing Liu, Ke WangAAAI 2026
- SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image CaptioningSi-Woo Kim, MinJu Jeon, Ye-Chan Kim, Soeun Lee et al.ACM MM 2025
- NLIP: Noise-Robust Language-Image Pre-trainingRunhui Huang, Yanxin Long, Jianhua Han, Hang Xu et al.AAAI 2023 · 44 citations
- NOC-REK: Novel Object Captioning with Retrieved Vocabulary from External KnowledgeDuc Minh Vo, Hong Chen, Akihiro Sugimoto, Hideki NakayamaCVPR 2022 · 21 citations
