Lune

NeurIPS2025Top-tier venue

Understanding the Gain from Data Filtering in Multimodal Contrastive Learning

Divyansh Pareek, Sewoong Oh, Simon S. Du

2025Year
1Citations
1Top-tier citations

Abstract

The success of modern multimodal representation learning relies on internet-scale datasets. Due to the low quality of a large fraction of raw web data, data curation has become a critical step in the training pipeline. Filtering using a trained model (i.e., teacher-based filtering) has emerged as a successful solution, leveraging a pre-trained model to compute quality scores. To explain the empirical success of teacher-based filtering, we characterize the performance of filtered contrastive learning under the standard bimodal data generation model. Denoting η∈(0,1]\eta\in(0,1] as the fraction of data with correctly matched modalities among nn paired samples, we utilize a linear contrastive learning setup to show a provable benefit of data filtering: (i)(i) the error without filtering is upper and lower bounded by 1ηn\frac{1}{\eta \sqrt{n}}, and (ii)(ii) the error with teacher-based filtering is upper bounded by 1ηn\frac{1}{\sqrt{\eta n}} in the large η\eta regime, and by 1n\frac{1}{\sqrt{n}} in the small η\eta regime.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 1f2ba711-537d-484d-99ef-c5bfe25794fe

Cited by top-tier papers1

Ask how each one uses it

Builds on21

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines