CovMatch: Cross-Covariance Guided Multimodal Dataset Distillation with Trainable Text Encoder
Yongmin Lee, Hye Won Chung
Abstract
Multimodal dataset distillation aims to synthesize a small set of image-text pairs that enables efficient training of large-scale vision-language models. While dataset distillation has shown promise in unimodal tasks, extending it to multimodal contrastive learning presents key challenges: learning cross-modal alignment and managing the high computational cost of large encoders. Prior approaches address scalability by freezing the text encoder and updating only the image encoder and text projection layer. However, we find this severely limits semantic alignment and becomes a bottleneck for performance scaling. We propose Cov-Match, a scalable dataset distillation framework that aligns the cross-covariances of real and synthetic features while regularizing feature distributions within each modality. Unlike prior approaches, CovMatch enables joint optimization of both encoders, leading to stronger cross-modal alignment and improved performance. Evaluated on Flickr30K and COCO, CovMatch outperforms state-ofthe-art multimodal distillation methods and achieves up to 6.8% absolute gains in retrieval accuracy using only 500 synthetic pairs. Our code is available at https://github.com/Yongalls/CovMatch.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7bdcc5e3-2f50-41b7-9a23-ba612f7f5490Cited by top-tier papers1
Ask how each one uses itBuilds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
Related papers
- Multimodal Distribution Matching for Vision-Language Dataset DistillationJongoh Jeong, Hoyong Kwon, Minseok Kim, Kuk-Jin YoonCVPR 2026 · 3 citations
- Efficient Multimodal Dataset Distillation via Generative ModelsZhenghao Zhao, Haoxuan Wang, Junyi Wu, Yuzhang Shang et al.NeurIPS 2025 · 7 citations
- Beyond Modality Collapse: Representation Blending for Multimodal Dataset DistillationXin Zhang, Ziruo Zhang, Jiawei Du, Zuozhu Liu et al.NeurIPS 2025 · 9 citations
- Asynchronous Matching with Dynamic Sampling for Multimodal Dataset DistillationDing Qi, Jian Li, Shuguang Dou, Zifan Song et al.ICLR 2026
- Low-Rank Similarity Mining for Multimodal Dataset DistillationYue Xu, Zhilin Lin, Yusong Qiu, Cewu Lu et al.ICML 2024 · 14 citations
