Filter Like You Test: Data-Driven Data Filtering for CLIP Pretraining
Mikey Shechter, Yair Carmon
摘要
We introduce Filter Like You Test (FLYT), an algorithm for curating large-scale vision-language datasets that learns the usefulness of each data point as a pretraining example. FLYT trains a scoring model that learns to weigh each example's features using gradient signals from downstream tasks training sets. Based on FLYT, we implement Mixing-FLYT (M-FLYT), which takes the per-example scores generated by different scoring methods as features, and learns to unify them into a single score. FLYT naturally produces a distribution over the training examples, which we leverage through Soft Cap Sampling (SCS), a strategy for obtaining a filtered pretraining dataset from per-example probabilities that samples examples while preventing over-representation through a repetition penalty. Using these methods, we achieve 40.1% ImageNet zero-shot accuracy on the DataComp medium scale filtering benchmark, a 2% absolute accuracy increase over all previous results and a 5.5% increase over results that - like us - use only public resources. Our approach also yields 37.7% on the average of 38 DataComp evaluation tasks, outperforming previous public-resource approaches by 0.4%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Understanding the Gain from Data Filtering in Multimodal Contrastive LearningDivyansh Pareek, Sewoong Oh, Simon S. DuNeurIPS 2025 · 被引用 1 次
- Matched Data, Better Models: Target Aligned Data Filtering with Sparse AutoencodersArnav Das, Gantavya Bhatt, Sahil Verma, Yiping Wang 等ICLR 2026
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli 等NeurIPS 2022 · 被引用 720 次
- Data Filtering NetworksAlex Fang, Albin Madappally Jose, Amit Jain, Ludwig Schmidt 等ICLR 2024 · 被引用 251 次
- Demystifying CLIP DataHu Xu, Saining Xie, Xiaoqing Ellen Tan, Po-Yao Huang 等ICLR 2024 · 被引用 249 次
- Characterizing Datapoints via Second-Split ForgettingPratyush Maini, Saurabh Garg, Zachary C. Lipton, J. Zico KolterNeurIPS 2022 · 被引用 48 次
相关 Paper
- Filtering, Distillation, and Hard Negatives for Vision-Language Pre-TrainingFilip Radenovic, Abhimanyu Dubey, Abhishek Kadian, Todor Mihaylov 等CVPR 2023
- Revisiting the Role of Language Priors in Vision-Language ModelsZhiqiu Lin, Xinyue Chen, Deepak Pathak, Pengchuan Zhang 等ICML 2024 · 被引用 44 次
- Few-Shot Recognition via Stage-Wise Retrieval-Augmented FinetuningTian Liu, Huixin Zhang, Shubham Parashar, Shu KongCVPR 2025
- Test-Time Prompt Tuning for Zero-Shot Generalization in Vision-Language ModelsManli Shu, Weili Nie, De-An Huang, Zhiding Yu 等NeurIPS 2022 · 被引用 603 次
- Prefix Conditioning Unifies Language and Label SupervisionKuniaki Saito, Kihyuk Sohn, Xiang Zhang, Chun-Liang Li 等CVPR 2023
