Understanding the Gain from Data Filtering in Multimodal Contrastive Learning
Divyansh Pareek, Sewoong Oh, Simon S. Du
Abstract
The success of modern multimodal representation learning relies on internet-scale datasets. Due to the low quality of a large fraction of raw web data, data curation has become a critical step in the training pipeline. Filtering using a trained model (i.e., teacher-based filtering) has emerged as a successful solution, leveraging a pre-trained model to compute quality scores. To explain the empirical success of teacher-based filtering, we characterize the performance of filtered contrastive learning under the standard bimodal data generation model. Denoting as the fraction of data with correctly matched modalities among paired samples, we utilize a linear contrastive learning setup to show a provable benefit of data filtering: the error without filtering is upper and lower bounded by , and the error with teacher-based filtering is upper bounded by in the large regime, and by in the small regime.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f2ba711-537d-484d-99ef-c5bfe25794feCited by top-tier papers1
Ask how each one uses itBuilds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Understanding Dimensional Collapse in Contrastive Self-supervised LearningLi Jing, Pascal Vincent, Yann LeCun, Yuandong TianICLR 2022 · 467 citations
- Data Selection for Language Models via Importance ResamplingSang Michael Xie, Shibani Santurkar, Tengyu Ma, Percy LiangNeurIPS 2023 · 383 citations
- Data Filtering NetworksAlex Fang, Albin Madappally Jose, Amit Jain, Ludwig Schmidt et al.ICLR 2024 · 251 citations
- Demystifying CLIP DataHu Xu, Saining Xie, Xiaoqing Ellen Tan, Po-Yao Huang et al.ICLR 2024 · 249 citations
Related papers
- Active Data Curation Effectively Distills Large-Scale Multimodal ModelsVishaal Udandarao, Nikhil Parthasarathy, Muhammad Ferjad Naeem, Talfan Evans et al.CVPR 2025
- LEMoN: Label Error Detection using Multimodal NeighborsHaoran Zhang, Aparna Balagopalan, Nassim Oufattole, Hyewon Jeong et al.ICML 2025
- CiT: Curation in Training for Effective Vision-Language DataHu Xu, Saining Xie, Po-Yao Huang, Licheng Yu et al.ICCV 2023 · 31 citations
- Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variablesYu Gui, Cong Ma, Zongming MaNeurIPS 2025 · 9 citations
- Code Representation Learning at ScaleDejiao Zhang, Wasi Uddin Ahmad, Ming Tan, Hantian Ding et al.ICLR 2024 · 32 citations
