Low-Rank Similarity Mining for Multimodal Dataset Distillation
Yue Xu, Zhilin Lin, Yusong Qiu, Cewu Lu, Yong-Lu Li
Abstract
Though dataset distillation has witnessed rapid development in recent years, the distillation of multimodal data, e.g., image-text pairs, poses unique and under-explored challenges. Unlike unimodal data, image-text contrastive learning (ITC) data lack inherent categorization and should instead place greater emphasis on modality correspondence. In this work, we propose Low-Rank Similarity Mining (LoRS) for multimodal dataset distillation, that concurrently distills a ground truth similarity matrix with image-text pairs, and leverages low-rank factorization for efficiency and scalability. The proposed approach brings significant improvement to the existing algorithms, marking a significant contribution to the field of visual-language dataset distillation. We advocate adopting LoRS as a foundational synthetic data setup for image-text dataset distillation. Our code is available at https: //github.com/silicx/LoRS_Distill .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ba98dacd-7bcd-4eaf-8cd4-8f022537f013Cited by top-tier papers10
- Beyond Modality Collapse: Representation Blending for Multimodal Dataset DistillationXin Zhang, Ziruo Zhang, Jiawei Du, Zuozhu Liu et al.NeurIPS 2025 · 9 citations
- Efficient Multimodal Dataset Distillation via Generative ModelsZhenghao Zhao, Haoxuan Wang, Junyi Wu, Yuzhang Shang et al.NeurIPS 2025 · 7 citations
- ImageBindDC: Compressing Multi-modal Data with ImageBind-based CondensationYue Min, Shaobo Wang, Jiaze Li, Tianle Niu et al.AAAI 2026 · 4 citations
- CovMatch: Cross-Covariance Guided Multimodal Dataset Distillation with Trainable Text EncoderYongmin Lee, Hye Won ChungNeurIPS 2025 · 2 citations
- Multimodal Dataset Distillation via Phased Teacher ModelsShengbin Guo, Hang Zhao, Senqiao Yang, Chenyang Jiang et al.ICLR 2026 · 1 citation
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
Related papers
- Multimodal Distribution Matching for Vision-Language Dataset DistillationJongoh Jeong, Hoyong Kwon, Minseok Kim, Kuk-Jin YoonCVPR 2026 · 3 citations
- Asynchronous Matching with Dynamic Sampling for Multimodal Dataset DistillationDing Qi, Jian Li, Shuguang Dou, Zifan Song et al.ICLR 2026
- Lightweight Contrastive Distilled Hashing for Online Cross-modal RetrievalJiaxing Li, Lin Jiang, Zeqi Ma, Kaihang Jiang et al.AAAI 2025 · 4 citations
- How to Make Cross Encoder a Good Teacher for Efficient Image-Text Retrieval?Yuxin Chen, Zongyang Ma, Ziqi Zhang, Zhongang Qi et al.CVPR 2024
- Multimodal Dataset Distillation Made Simple by Prototype-Guided Data SynthesisJunhyeok Choi, Sangwoo Mo, Minwoo ChaeICLR 2026
