Lune

NeurIPS2025Top-tier venue

Meta CLIP 2: A Worldwide Scaling Recipe

Yung-Sung Chuang, Yang Li, Dong Wang, Ching-Feng Yeh, Kehan Lyu, Ramya Raghavendra, Jim Glass, Lifei Huang, Jason E. Weston, Luke Zettlemoyer, Xinlei Chen, Zhuang Liu

2025Year
72Citations
17Top-tier citations

Abstract

Contrastive Language-Image Pretraining (CLIP) is a popular foundation model, supporting from zero-shot classification, retrieval to encoders for multimodal large language models (MLLMs). Although CLIP is successfully trained on billion-scale image-text pairs from the English world, scaling CLIP's training further to learning from the worldwide web data is still challenging: (1) no curation method is available to handle data points from non-English world; (2) the English performance from existing multilingual CLIP is worse than its English-only counterpart, i.e., "curse of multilinguality" that is common in LLMs. Here, we present Meta CLIP 2, the first recipe training CLIP from scratch on worldwide web-scale image-text pairs. To generalize our findings, we conduct rigorous ablations with minimal changes that are necessary to address the above challenges and present a recipe enabling mutual benefits from English and non-English world data. In zero-shot ImageNet classification, Meta CLIP 2 ViT-H/14 surpasses its English-only counterpart by 0.8% and mSigLIP by 0.7%, and surprisingly sets new state-of-the-art without system-level confounding factors (e.g., translation, bespoke architecture changes) on multilingual benchmarks, such as CVQA with 57.4%, Babel-ImageNet with 50.2% and XM3600 with 64.3% on image-to-text retrieval. Code and model are available at https://github.com/facebookresearch/MetaCLIP. breaks the curse. English accuracy rises from 80.5% to 81.3% on ImageNet and surprisingly new SoTA is set with minimal CLIP architecture changes for multilingual image-to-text retrieval (XM3600 64.3%, Babel-ImageNet 50.2%, and CVQA 57.4%).

Together, Meta CLIP 2 enables the following desirable results by nature. 1) Mutual benefits from the English and non-English worlds. Non-English data now can better support an Englishonly model and vice versa, which is critical in the era when English data is depleting. 2) Full multilingual support. Meta CLIP 2 never drops image-text pairs simply by languages and yields models outperforming all the previous multilingual systems, such as mSigLIP [16] and SigLIP 2 [17]. 3) Native-language supervision. Models learn directly from alt-texts written by native speakers rather than synthetic machine translations [21,14]. 4) Cultural diversity. Meta CLIP 2 retains the entire global distribution of images and thus inherits the comprehensive cultural and socioeconomic coverage advocated by [21]. Such coverage improves geo-localization and region-specific recognition. 5) No-filter philosophy. With the curation algorithm designed towards worldwide data, Meta CLIP 2 removes the last filter (i.e., whether the alt-text is in English) in pipeline, achieving better diversity and minimizing biases introduced by filters [21]. 6) Broader impacts on foundation data. This work provides a foundational data algorithm designed for worldwide scale, and benefits not only CLIP, but also efforts using CLIP data such as MLLM [2, 22], SSL (Web-DINO [23]) and image generation and diffusion models [25]).

2 Related Work 2.1 Evolution of CLIP and its Data Processing CLIP [1] and its variants [26,6,16] learn versatile image and text representations that are generally useful for downstream tasks [2,27,4]. Such multimodal contrastive learning and transformer architectures become standard components in vision and multimodal research. Data is a key contributor to CLIP's performance [28,7]. Two major processing approaches for CLIP data emerge: curation 5 from scratch, and distillation from external resources. One key difference is that the former yields more controllable distribution and the latter has intractable distribution owned by an outsourcing party.

Curation from scratch. OpenAI CLIP [1] curates a training dataset of 400M image-text pairs from scratch and publicizes high-level curation guidance. Meta CLIP [7] makes OpenAI's guidance as a formal curation algorithm and scales the curation to 2.5B pairs. The algorithm is model-free, no blackbox filtering, and fully transparent to enable training entirely from scratch on public data source, where the data distribution is curated to align with metadata composed by human experts (e.g., WordNet and Wikipedia).

Distillation from external resources. Distillation-based methods usually have good performance and save compute by learning from teacher model's knowledge [30]. However, in the context of CLIP training the teacher is usually an external blackbox system, which introduces intractable bias. For example, LAION-400M/5B [8,31] (used by OpenCLIP [6]) relies on OpenAI CLIP-filter and DFN [10] using a filter model trained on high-quality private data [32]. Recently, SigLIP [16] and SigLIP 2 [17] learn from data source WebLI [15], which is derived from Google Image Search [18].

CLIP-style models are widely used as vision encoders in MLLM, where language supervision in CLIP training helps to learn compact and semantic-rich visual representations. In contrast, traditiona

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 59a3f4cc-a4c9-4e03-a0f2-6ef3529c080e

Cited by top-tier papers17

Ask how each one uses it

Builds on26

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines