Lune

NeurIPS2025顶会

Meta CLIP 2: A Worldwide Scaling Recipe

Yung-Sung Chuang, Yang Li, Dong Wang, Ching-Feng Yeh, Kehan Lyu, Ramya Raghavendra, Jim Glass, Lifei Huang, Jason E. Weston, Luke Zettlemoyer, Xinlei Chen, Zhuang Liu

2025年份
72被引次数
17顶会引用

摘要

Contrastive Language-Image Pretraining (CLIP) is a popular foundation model, supporting from zero-shot classification, retrieval to encoders for multimodal large language models (MLLMs). Although CLIP is successfully trained on billion-scale image-text pairs from the English world, scaling CLIP's training further to learning from the worldwide web data is still challenging: (1) no curation method is available to handle data points from non-English world; (2) the English performance from existing multilingual CLIP is worse than its English-only counterpart, i.e., "curse of multilinguality" that is common in LLMs. Here, we present Meta CLIP 2, the first recipe training CLIP from scratch on worldwide web-scale image-text pairs. To generalize our findings, we conduct rigorous ablations with minimal changes that are necessary to address the above challenges and present a recipe enabling mutual benefits from English and non-English world data. In zero-shot ImageNet classification, Meta CLIP 2 ViT-H/14 surpasses its English-only counterpart by 0.8% and mSigLIP by 0.7%, and surprisingly sets new state-of-the-art without system-level confounding factors (e.g., translation, bespoke architecture changes) on multilingual benchmarks, such as CVQA with 57.4%, Babel-ImageNet with 50.2% and XM3600 with 64.3% on image-to-text retrieval. Code and model are available at https://github.com/facebookresearch/MetaCLIP. breaks the curse. English accuracy rises from 80.5% to 81.3% on ImageNet and surprisingly new SoTA is set with minimal CLIP architecture changes for multilingual image-to-text retrieval (XM3600 64.3%, Babel-ImageNet 50.2%, and CVQA 57.4%).

Together, Meta CLIP 2 enables the following desirable results by nature. 1) Mutual benefits from the English and non-English worlds. Non-English data now can better support an Englishonly model and vice versa, which is critical in the era when English data is depleting. 2) Full multilingual support. Meta CLIP 2 never drops image-text pairs simply by languages and yields models outperforming all the previous multilingual systems, such as mSigLIP [16] and SigLIP 2 [17]. 3) Native-language supervision. Models learn directly from alt-texts written by native speakers rather than synthetic machine translations [21,14]. 4) Cultural diversity. Meta CLIP 2 retains the entire global distribution of images and thus inherits the comprehensive cultural and socioeconomic coverage advocated by [21]. Such coverage improves geo-localization and region-specific recognition. 5) No-filter philosophy. With the curation algorithm designed towards worldwide data, Meta CLIP 2 removes the last filter (i.e., whether the alt-text is in English) in pipeline, achieving better diversity and minimizing biases introduced by filters [21]. 6) Broader impacts on foundation data. This work provides a foundational data algorithm designed for worldwide scale, and benefits not only CLIP, but also efforts using CLIP data such as MLLM [2, 22], SSL (Web-DINO [23]) and image generation and diffusion models [25]).

2 Related Work 2.1 Evolution of CLIP and its Data Processing CLIP [1] and its variants [26,6,16] learn versatile image and text representations that are generally useful for downstream tasks [2,27,4]. Such multimodal contrastive learning and transformer architectures become standard components in vision and multimodal research. Data is a key contributor to CLIP's performance [28,7]. Two major processing approaches for CLIP data emerge: curation 5 from scratch, and distillation from external resources. One key difference is that the former yields more controllable distribution and the latter has intractable distribution owned by an outsourcing party.

Curation from scratch. OpenAI CLIP [1] curates a training dataset of 400M image-text pairs from scratch and publicizes high-level curation guidance. Meta CLIP [7] makes OpenAI's guidance as a formal curation algorithm and scales the curation to 2.5B pairs. The algorithm is model-free, no blackbox filtering, and fully transparent to enable training entirely from scratch on public data source, where the data distribution is curated to align with metadata composed by human experts (e.g., WordNet and Wikipedia).

Distillation from external resources. Distillation-based methods usually have good performance and save compute by learning from teacher model's knowledge [30]. However, in the context of CLIP training the teacher is usually an external blackbox system, which introduces intractable bias. For example, LAION-400M/5B [8,31] (used by OpenCLIP [6]) relies on OpenAI CLIP-filter and DFN [10] using a filter model trained on high-quality private data [32]. Recently, SigLIP [16] and SigLIP 2 [17] learn from data source WebLI [15], which is derived from Google Image Search [18].

CLIP-style models are widely used as vision encoders in MLLM, where language supervision in CLIP training helps to learn compact and semantic-rich visual representations. In contrast, traditiona

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 59a3f4cc-a4c9-4e03-a0f2-6ef3529c080e

引用它的顶会 Paper17

问问它们各自怎么用它

它引用的顶会 Paper26

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖