Meta CLIP 2: A Worldwide Scaling Recipe
Yung-Sung Chuang, Yang Li, Dong Wang, Ching-Feng Yeh, Kehan Lyu, Ramya Raghavendra, Jim Glass, Lifei Huang, Jason E. Weston, Luke Zettlemoyer, Xinlei Chen, Zhuang Liu
摘要
Contrastive Language-Image Pretraining (CLIP) is a popular foundation model, supporting from zero-shot classification, retrieval to encoders for multimodal large language models (MLLMs). Although CLIP is successfully trained on billion-scale image-text pairs from the English world, scaling CLIP's training further to learning from the worldwide web data is still challenging: (1) no curation method is available to handle data points from non-English world; (2) the English performance from existing multilingual CLIP is worse than its English-only counterpart, i.e., "curse of multilinguality" that is common in LLMs. Here, we present Meta CLIP 2, the first recipe training CLIP from scratch on worldwide web-scale image-text pairs. To generalize our findings, we conduct rigorous ablations with minimal changes that are necessary to address the above challenges and present a recipe enabling mutual benefits from English and non-English world data. In zero-shot ImageNet classification, Meta CLIP 2 ViT-H/14 surpasses its English-only counterpart by 0.8% and mSigLIP by 0.7%, and surprisingly sets new state-of-the-art without system-level confounding factors (e.g., translation, bespoke architecture changes) on multilingual benchmarks, such as CVQA with 57.4%, Babel-ImageNet with 50.2% and XM3600 with 64.3% on image-to-text retrieval. Code and model are available at https://github.com/facebookresearch/MetaCLIP. breaks the curse. English accuracy rises from 80.5% to 81.3% on ImageNet and surprisingly new SoTA is set with minimal CLIP architecture changes for multilingual image-to-text retrieval (XM3600 64.3%, Babel-ImageNet 50.2%, and CVQA 57.4%).
Together, Meta CLIP 2 enables the following desirable results by nature. 1) Mutual benefits from the English and non-English worlds. Non-English data now can better support an Englishonly model and vice versa, which is critical in the era when English data is depleting. 2) Full multilingual support. Meta CLIP 2 never drops image-text pairs simply by languages and yields models outperforming all the previous multilingual systems, such as mSigLIP [16] and SigLIP 2 [17]. 3) Native-language supervision. Models learn directly from alt-texts written by native speakers rather than synthetic machine translations [21,14]. 4) Cultural diversity. Meta CLIP 2 retains the entire global distribution of images and thus inherits the comprehensive cultural and socioeconomic coverage advocated by [21]. Such coverage improves geo-localization and region-specific recognition. 5) No-filter philosophy. With the curation algorithm designed towards worldwide data, Meta CLIP 2 removes the last filter (i.e., whether the alt-text is in English) in pipeline, achieving better diversity and minimizing biases introduced by filters [21]. 6) Broader impacts on foundation data. This work provides a foundational data algorithm designed for worldwide scale, and benefits not only CLIP, but also efforts using CLIP data such as MLLM [2, 22], SSL (Web-DINO [23]) and image generation and diffusion models [25]).
2 Related Work 2.1 Evolution of CLIP and its Data Processing CLIP [1] and its variants [26,6,16] learn versatile image and text representations that are generally useful for downstream tasks [2,27,4]. Such multimodal contrastive learning and transformer architectures become standard components in vision and multimodal research. Data is a key contributor to CLIP's performance [28,7]. Two major processing approaches for CLIP data emerge: curation 5 from scratch, and distillation from external resources. One key difference is that the former yields more controllable distribution and the latter has intractable distribution owned by an outsourcing party.
Curation from scratch. OpenAI CLIP [1] curates a training dataset of 400M image-text pairs from scratch and publicizes high-level curation guidance. Meta CLIP [7] makes OpenAI's guidance as a formal curation algorithm and scales the curation to 2.5B pairs. The algorithm is model-free, no blackbox filtering, and fully transparent to enable training entirely from scratch on public data source, where the data distribution is curated to align with metadata composed by human experts (e.g., WordNet and Wikipedia).
Distillation from external resources. Distillation-based methods usually have good performance and save compute by learning from teacher model's knowledge [30]. However, in the context of CLIP training the teacher is usually an external blackbox system, which introduces intractable bias. For example, LAION-400M/5B [8,31] (used by OpenCLIP [6]) relies on OpenAI CLIP-filter and DFN [10] using a filter model trained on high-quality private data [32]. Recently, SigLIP [16] and SigLIP 2 [17] learn from data source WebLI [15], which is derived from Google Image Search [18].
CLIP-style models are widely used as vision encoders in MLLM, where language supervision in CLIP training helps to learn compact and semantic-rich visual representations. In contrast, traditiona
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late InteractionZilin Xiao, Qi Ma, Mengting Gu, Chun-cheng Jason Chen 等ICLR 2026 · 被引用 40 次
- Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal PerceptionLai Wei, Liangbo He, jun lan, Lingzhong Dong 等ICML 2026 · 被引用 27 次
- Franca: Nested Matryoshka Clustering for Scalable Visual Representation LearningShashanka Venkataramanan, Valentinos Pariza, Mohammadreza Salehi, Lukas Knobel 等CVPR 2026 · 被引用 26 次
- In Pursuit of Pixel Supervision for Visual Pre-trainingLihe Yang, Shang-Wen Li, Yang Li, Xinjie Lei 等CVPR 2026 · 被引用 13 次
- The Information Geometry of Softmax: Probing and SteeringKiho Park, Todd Nief, Yo Joong Choe, Victor VeitchICML 2026 · 被引用 6 次
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
相关 Paper
- Demystifying CLIP DataHu Xu, Saining Xie, Xiaoqing Ellen Tan, Po-Yao Huang 等ICLR 2024 · 被引用 249 次
- mCLIP: Multilingual CLIP via Cross-lingual TransferGuanhua Chen, Lu Hou, Yun Chen, Wenliang Dai 等ACL 2023 · 被引用 13 次
- Building Vision-Language Models on Solid Foundations with Masked DistillationSepehr Sameni, Kushal Kafle, Hao Tan, Simon JenniCVPR 2024 · 被引用 4 次
- FG-CLIP: Fine-Grained Visual and Textual AlignmentChunyu Xie, Bin Wang, Fanjing Kong, Jincheng Li 等ICML 2025
- LLM2CLIP: Powerful Language Model Unlocks Richer Cross-Modality RepresentationWeiquan Huang, Aoqi Wu, Yifan Yang, Xufang Luo 等AAAI 2026
