FALCON: False-Negative Aware Learning of Contrastive Negatives in Vision-Language Alignment
Myunsoo Kim, Seong-Woong Shim, Byung-Jun Lee
Abstract
False negatives pose a critical challenge in vision-language pretraining (VLP) due to the many-to-many correspondence between images and texts in large-scale datasets. These false negatives introduce conflicting supervision signals that degrade the learned embedding space and diminish the effectiveness of hard negative sampling. In this paper, we propose FALCON (False-negative Aware Learning of COntrastive Negatives), a learning-based mini-batch construction strategy that adaptively balances the trade-off between hard and false negatives during VLP. Rather than relying on fixed heuristics, FALCON employs a negative mining scheduler that dynamically selects negative samples of appropriate hardness for each anchor instance during mini-batch construction, guided by a proxy for cross-modal alignment improvement. Experimental results demonstrate that FALCON significantly improves performance across three vision-language learning frameworks (ALBEF, BLIP-2, SigLIP-2) and a broad range of downstream tasks and evaluation settings, underscoring its effectiveness and robustness in mitigating the impact of false negatives.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd1fa2a2-e17f-44d8-90fc-20050be34af0Cited by top-tier papers1
Ask how each one uses itBuilds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
Related papers
- MAFA: Managing False Negatives for Vision-Language Pre-TrainingJaeseok Byun, Dohoon Kim, Taesup MoonCVPR 2024
- FaNe: Towards Fine-Grained Cross-Modal Contrast with False-Negative Reduction and Text-Conditioned Sparse AttentionPeng Zhang, Zhihui Lai, Wenting Chen, Xu Wu et al.AAAI 2026
- Selective Contrastive Learning For Gloss Free Sign Language TranslationChanghao Lai, Rui Zhao, Xuewen Zhong, Jinsong Su et al.ACL 2026
- Importance Sampling for Multi-Negative Multimodal Direct Preference OptimizationXintong Li, Chuhan Wang, Junda Wu, Rohan Surana et al.ICLR 2026 · 5 citations
- M2-VLP: Enhancing Multilingual Vision-Language Pre-Training via Multi-Grained AlignmentAhtamjan Ahmat, Lei Wang, Yating Yang, Bo Ma et al.WWW 2025 · 2 citations
