FILIP: Fine-grained Interactive Language-Image Pre-Training
Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, Chunjing Xu
摘要
Unsupervised large-scale vision-language pre-training has shown promising advances on various downstream tasks. Existing methods often model the cross-modal interaction either via the similarity of the global feature of each modality which misses sufficient information, or finer-grained interactions using cross/self-attention upon visual and textual tokens. However, cross/self-attention suffers from inferior efficiency in both training and inference. In this paper, we introduce a large-scale Fine-grained Interactive Language-Image Pre-training (FILIP) to achieve finer-level alignment through a cross-modal late interaction mechanism, which uses a token-wise maximum similarity between visual and textual tokens to guide the contrastive objective. FILIP successfully leverages the finer-grained expressiveness between image patches and textual words by modifying only contrastive loss, while simultaneously gaining the ability to pre-compute image and text representations offline at inference, keeping both large-scale training and inference efficient. Furthermore, we construct a new large-scale image-text pair dataset called FILIP300M for pre-training. Experiments show that FILIP achieves state-of-the-art performance on multiple downstream vision-language tasks including zero-shot image classification and image-text retrieval. The visualization on word-patch alignment further shows that FILIP can learn meaningful fine-grained features with promising localization ability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper320
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training ParadigmYangguang Li, Feng Liang, Lichen Zhao, Yufeng Cui 等ICLR 2022 · 被引用 565 次
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li 等ICLR 2024 · 被引用 467 次
- Self-regulating Prompts: Foundational Model Adaptation without ForgettingMuhammad Uzair Khattak, Syed Talal Wasim, Muzammal Naseer, Salman Khan 等ICCV 2023 · 被引用 365 次
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
相关 Paper
- FineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text InputsMothilal Asokan, Kebin Wu, Fatima AlbreikiCVPR 2025
- HiVLP: Hierarchical Interactive Video-Language Pre-TrainingBin Shao, Jianzhuang Liu, Renjing Pei, Songcen Xu 等ICCV 2023 · 被引用 6 次
- PyramidCLIP: Hierarchical Feature Alignment for Vision-language Model PretrainingYuting Gao, Jinfeng Liu, Zihan Xu, Jun Zhang 等NeurIPS 2022 · 被引用 168 次
- CALIP: Zero-Shot Enhancement of CLIP with Parameter-Free AttentionZiyu Guo, Renrui Zhang, Longtian Qiu, Xianzheng Ma 等AAAI 2023 · 被引用 182 次
- M2-VLP: Enhancing Multilingual Vision-Language Pre-Training via Multi-Grained AlignmentAhtamjan Ahmat, Lei Wang, Yating Yang, Bo Ma 等WWW 2025 · 被引用 2 次
