Too Large; Data Reduction for Vision-Language Pre-Training
Alex Jinpeng Wang, Kevin Qinghong Lin, David Junhao Zhang, Stan Weixian Lei, Mike Zheng Shou
Abstract
This paper examines the problems of severe image-text misalignment and high redundancy in the widely-used large-scale Vision-Language Pre-Training (VLP) datasets. To address these issues, we propose an efficient and straightforward Vision-Language learning algorithm called , which aims to compress the existing large VLP data into a small, high-quality set. Our approach consists of two major steps. First, a codebook-based encoder-decoder captioner is developed to select representative samples. Second, a new caption is generated to complement the original captions for selected samples, mitigating the text-image misalignment problem while maintaining uniqueness. As the result, enables us to reduce the large dataset into a small set of high-quality data, which can serve as an alternative pre-training dataset. This algorithm significantly speeds up the time-consuming pretraining process. Specifically, can compress the mainstream VLP datasets at a high ratio, e.g., reduce well-cleaned CC3M dataset from 2.82M to 0.67M ( 24%) and noisy YFCC15M from 15M to 2.5M ( 16.7%). Extensive experiments with three popular VLP models over seven downstream tasks show that VLP model trained on the compressed dataset provided by can perform similar or even better results compared with training on the full-scale dataset1.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e34c10a-867a-4301-8960-a3fca9b312d9Cited by top-tier papers12
- SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-ImprovementXiyao Wang, Zhengyuan Yang, Chao Feng, Hongjin Lu et al.NeurIPS 2025 · 158 citations
- Effective pruning of web-scale datasets based on complexity of concept clustersAmro Abbas, Evgenia Rusak, Kushal Tirumala, Wieland Brendel et al.ICLR 2024 · 30 citations
- Data-Efficient Multimodal Fusion on a Single GPUNoël Vouitsis, Zhaoyan Liu, Satya Krishna Gorti, Valentin Villecroze et al.CVPR 2024 · 6 citations
- SaCo Loss: Sample-Wise Affinity Consistency for Vision-Language Pre-TrainingSitong Wu, Haoru Tan, Zhuotao Tian, Yukang Chen et al.CVPR 2024 · 5 citations
- Mixture-of-Scores: Robust Image-Text Data Valuation via Three Lines of CodeSitong Wu, Haoru Tan, Yukang Chen, Shaofeng Zhang et al.ICCV 2025 · 4 citations
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Double-Filter: Efficient Fine-tuning of Pre-trained Vision-Language Models via Patch&Layer FilteringYaoqin He, Junchen Fu, Kaiwen Zheng, Songpei Xu et al.ICML 2025
- Expedited Training of Visual Conditioned Language Generation via Redundancy ReductionYiren Jian, Tingkai Liu, Yunzhe Tao, Chunhui Zhang et al.ACL 2024 · 5 citations
- Accelerating Vision-Language Pretraining with Free Language ModelingTeng Wang, Yixiao Ge, Feng Zheng, Ran Cheng et al.CVPR 2023
- CapEnrich: Enriching Caption Semantics for Web Images via Cross-modal Pre-trained KnowledgeLinli Yao, Weijing Chen, Qin JinWWW 2023 · 11 citations
- Parameter and Computation Efficient Transfer Learning for Vision-Language Pre-trained ModelsQiong Wu, Wei Yu, Yiyi Zhou, Shubin Huang et al.NeurIPS 2023 · 16 citations
