MAFA: Managing False Negatives for Vision-Language Pre-Training
Jaeseok Byun, Dohoon Kim, Taesup Moon
Abstract
We consider a critical issue of false negatives in Vision-Language Pre-training (VLP), a challenge that arises from the inherent many-to-many correspondence of image-text pairs in large-scale web-crawled datasets. The presence of false negatives can impede achieving optimal performance and even lead to a significant performance drop. To address this challenge, we propose MAFA (MAnaging FAlse negatives), which consists of two pivotal components building upon the recently developed GRouped mIni-baTch sampling (GRIT) strategy: 1) an efficient connection mining process that identifies and converts false negatives into positives, and 2) label smoothing for the image-text contrastive (ITC) loss. Our comprehensive experiments verify the effectiveness of MAFA across multiple downstream tasks, emphasizing the crucial role of addressing false negatives in VLP, potentially even surpassing the importance of addressing false positives. In addition, the compatibility of MAFA with the recent BLIP-family model is also demonstrated. Code is available at https://github.com/jaeseokbyun/MAFA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b8a43cbe-56b2-44b4-97a0-dbc82b93555eCited by top-tier papers4
- CellCLIP - Learning Perturbation Effects in Cell Painting via Text-Guided Contrastive LearningMingyu Lu, Ethan Weinberger, Chanwoo Kim, Su-In LeeNeurIPS 2025 · 9 citations
- FALCON: False-Negative Aware Learning of Contrastive Negatives in Vision-Language AlignmentMyunsoo Kim, Seong-Woong Shim, Byung-Jun LeeCVPR 2026 · 2 citations
- Concept-Aware Batch Sampling Improves Language-Image PretrainingAdhiraj Ghosh, Vishaal Udandarao, Thao Nguyen, Matteo Farina et al.CVPR 2026 · 1 citation
- GeoMM: On Geodesic Perspective for Multi-modal LearningShibin Mei, Hang Wang, Bingbing NiCVPR 2025
Builds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- ViLTA: Enhancing Vision-Language Pre-training through Textual AugmentationWeihan Wang, Zhen Yang, Bin Xu, Juanzi Li et al.ICCV 2023 · 11 citations
- FaNe: Towards Fine-Grained Cross-Modal Contrast with False-Negative Reduction and Text-Conditioned Sparse AttentionPeng Zhang, Zhihui Lai, Wenting Chen, Xu Wu et al.AAAI 2026
- FFF: Fixing Flawed Foundations in contrastive pre-training results in very strong Vision-Language modelsAdrian Bulat, Yassine Ouali, Georgios TzimiropoulosCVPR 2024
- ALIP: Adaptive Language-Image Pre-training with Synthetic CaptionKaicheng Yang, Jiankang Deng, Xiang An, Jiawei Li et al.ICCV 2023 · 93 citations
- How Much Can CLIP Benefit Vision-and-Language Tasks?Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal et al.ICLR 2022 · 503 citations
