FFF: Fixing Flawed Foundations in contrastive pre-training results in very strong Vision-Language models
Adrian Bulat, Yassine Ouali, Georgios Tzimiropoulos
Abstract
Despite noise and caption quality having been acknowledged as important factors impacting vision-language contrastive pre-training, in this paper, we show that the full potential of improving the training process by addressing such issues is yet to be realized. Specifically, we firstly study and analyze two issues affecting training: incorrect assignment of negative pairs, and low caption quality and diversity. Then, we devise effective solutions for addressing both problems, which essentially require training with multiple true positive pairs. Finally, we propose training with sigmoid loss to address such a requirement. We show very large gains over the current state-of-the-art for both image recognition (∼ +6% on average over 11 datasets) and image retrieval (∼ +19% on Flickr30k and ∼ +15% on MSCOCO).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dddd82c9-9700-44eb-ac55-61769c3ebcd2Cited by top-tier papers5
- Unlearning the Noisy Correspondence Makes CLIP More RobustHaochen Han, Alex Jinpeng Wang, Peijun Ye, Fangming LiuICCV 2025 · 3 citations
- FALCON: False-Negative Aware Learning of Contrastive Negatives in Vision-Language AlignmentMyunsoo Kim, Seong-Woong Shim, Byung-Jun LeeCVPR 2026 · 2 citations
- Efficient Vision-Language pre-training via domain-specific learning for human activitiesAdrian Bulat, Yassine Ouali, Ricardo Guerrero, Brais Martínez et al.EMNLP 2024 · 1 citation
- VladVA: Discriminative Fine-tuning of LVLMsYassine Ouali, Adrian Bulat, Alexandros Xenos, Anestis Zaganidis et al.CVPR 2025
- Enhancing Vision-Language Compositional Understanding with Multimodal Synthetic DataHaoxin Li, Boyang LiCVPR 2025
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- ALIP: Adaptive Language-Image Pre-training with Synthetic CaptionKaicheng Yang, Jiankang Deng, Xiang An, Jiawei Li et al.ICCV 2023 · 93 citations
- MAFA: Managing False Negatives for Vision-Language Pre-TrainingJaeseok Byun, Dohoon Kim, Taesup MoonCVPR 2024
- Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only TrainingLongtian Qiu, Shan Ning, Xuming HeAAAI 2024 · 20 citations
- TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language NegativesMaitreya Patel, Abhiram Kusumba, Sheng Cheng, Changhoon Kim et al.NeurIPS 2024 · 73 citations
- Robust Cross-Modal Representation Learning with Progressive Self-DistillationAlex Andonian, Shixing Chen, Raffay HamidCVPR 2022 · 43 citations
