Filter Like You Test: Data-Driven Data Filtering for CLIP Pretraining
Mikey Shechter, Yair Carmon
Abstract
We introduce Filter Like You Test (FLYT), an algorithm for curating large-scale vision-language datasets that learns the usefulness of each data point as a pretraining example. FLYT trains a scoring model that learns to weigh each example's features using gradient signals from downstream tasks training sets. Based on FLYT, we implement Mixing-FLYT (M-FLYT), which takes the per-example scores generated by different scoring methods as features, and learns to unify them into a single score. FLYT naturally produces a distribution over the training examples, which we leverage through Soft Cap Sampling (SCS), a strategy for obtaining a filtered pretraining dataset from per-example probabilities that samples examples while preventing over-representation through a repetition penalty. Using these methods, we achieve 40.1% ImageNet zero-shot accuracy on the DataComp medium scale filtering benchmark, a 2% absolute accuracy increase over all previous results and a 5.5% increase over results that - like us - use only public resources. Our approach also yields 37.7% on the average of 38 DataComp evaluation tasks, outperforming previous public-resource approaches by 0.4%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f941bc9-e759-4919-9540-26e7440e9cafCited by top-tier papers2
- Understanding the Gain from Data Filtering in Multimodal Contrastive LearningDivyansh Pareek, Sewoong Oh, Simon S. DuNeurIPS 2025 · 1 citation
- Matched Data, Better Models: Target Aligned Data Filtering with Sparse AutoencodersArnav Das, Gantavya Bhatt, Sahil Verma, Yiping Wang et al.ICLR 2026
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
- Data Filtering NetworksAlex Fang, Albin Madappally Jose, Amit Jain, Ludwig Schmidt et al.ICLR 2024 · 251 citations
- Demystifying CLIP DataHu Xu, Saining Xie, Xiaoqing Ellen Tan, Po-Yao Huang et al.ICLR 2024 · 249 citations
- Characterizing Datapoints via Second-Split ForgettingPratyush Maini, Saurabh Garg, Zachary C. Lipton, J. Zico KolterNeurIPS 2022 · 48 citations
Related papers
- Filtering, Distillation, and Hard Negatives for Vision-Language Pre-TrainingFilip Radenovic, Abhimanyu Dubey, Abhishek Kadian, Todor Mihaylov et al.CVPR 2023
- Revisiting the Role of Language Priors in Vision-Language ModelsZhiqiu Lin, Xinyue Chen, Deepak Pathak, Pengchuan Zhang et al.ICML 2024 · 44 citations
- Few-Shot Recognition via Stage-Wise Retrieval-Augmented FinetuningTian Liu, Huixin Zhang, Shubham Parashar, Shu KongCVPR 2025
- Test-Time Prompt Tuning for Zero-Shot Generalization in Vision-Language ModelsManli Shu, Weili Nie, De-An Huang, Zhiding Yu et al.NeurIPS 2022 · 603 citations
- Prefix Conditioning Unifies Language and Label SupervisionKuniaki Saito, Kihyuk Sohn, Xiang Zhang, Chun-Liang Li et al.CVPR 2023
