Noise-Aware Image Captioning with Progressively Exploring Mismatched Words
Zhongtian Fu, Kefei Song, Luping Zhou, Yang Yang
摘要
Image captioning aims to automatically generate captions for images by learning a cross-modal generator from vision to language. The large amount of image-text pairs required for training is usually sourced from the internet due to the manual cost, which brings the noise with mismatched relevance that affects the learning process. Unlike traditional noisy label learning, the key challenge in processing noisy image-text pairs is to finely identify the mismatched words to make the most use of trustworthy information in the text, rather than coarsely weighing the entire examples. To tackle this challenge, we propose a Noise-aware Image Captioning method (NIC) to adaptively mitigate the erroneous guidance from noise by progressively exploring mismatched words. Specifically, NIC first identifies mismatched words by quantifying word-label reliability from two aspects: 1) inter-modal representativeness, which measures the significance of the current word by assessing cross-modal correlation via prediction certainty; 2) intra-modal informativeness, which amplifies the effect of current prediction by combining the quality of subsequent word generation. During optimization, NIC constructs the pseudo-word-labels considering the reliability of the origin word-labels and model convergence to periodically coordinate mismatched words. As a result, NIC can effectively exploit both clean and noisy image-text pairs to learn a more robust mapping function. Extensive experiments conducted on the MS-COCO and Conceptual Caption datasets validate the effectiveness of our method in various noisy scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Facilitating Multimodal Classification via Dynamically Learning Modality GapYang Yang, Fengqiang Wan, Qing-Yuan Jiang, Yi XuNeurIPS 2024 · 被引用 65 次
- UICopilot: Automating UI Synthesis via Hierarchical Code Generation from Webpage DesignsYi Gui, Yao Wan, Zhen Li, Zhongyi Zhang 等WWW 2025 · 被引用 24 次
- ConeSep: Cone-based Robust Noise-Unlearning Compositional Network for Composed Image RetrievalZixu Li, Yupeng Hu, Zhiwei Chen, Mingyu Zhang 等CVPR 2026 · 被引用 16 次
- Unlearning the Noisy Correspondence Makes CLIP More RobustHaochen Han, Alex Jinpeng Wang, Peijun Ye, Fangming LiuICCV 2025 · 被引用 3 次
- Text as Any-Modality for Zero-Shot Classification by Consistent Prompt TuningXiangyu Wu, Feng Yu, Yang Yang, Jianfeng LuACM MM 2025
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- Symmetric Cross Entropy for Robust Learning With Noisy LabelsYisen Wang, Xingjun Ma, Zaiyi Chen, Yuan Luo 等ICCV 2019 · 被引用 1,125 次
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 被引用 992 次
- Learning with Noisy Correspondence for Cross-modal MatchingZhenyu Huang, Guocheng Niu, Xiao Liu, Wenbiao Ding 等NeurIPS 2021 · 被引用 215 次
相关 Paper
- Noise-aware Learning from Web-crawled Image-Text Data for Image CaptioningWooyoung Kang, Jonghwan Mun, Sungjun Lee, Byungseok RohICCV 2023 · 被引用 33 次
- PCSR: Pseudo-label Consistency-Guided Sample Refinement for Noisy Correspondence LearningZhuoyao Liu, Yang Liu, Wentao Feng, Shudong HuangAAAI 2026
- DECIDER: Difference-aware Contrastive Diffusion Model with Adversarial Perturbations for Image Change CaptioningGuojin Zhong, Jinhong Hu, Jiajun Chen, Jin Yuan 等AAAI 2025 · 被引用 3 次
- Show Your Faith: Cross-Modal Confidence-Aware Network for Image-Text MatchingHuatian Zhang, Zhendong Mao, Kun Zhang, Yongdong ZhangAAAI 2022 · 被引用 62 次
- PC2: Pseudo-Classification Based Pseudo-Captioning for Noisy Correspondence Learning in Cross-Modal RetrievalYue Duan, Zhangxuan Gu, Zhenzhe Ying, Lei Qi 等ACM MM 2024 · 被引用 10 次
