Diversified Semantic Distribution Matching for Dataset Distillation
Hongcheng Li, Yucan Zhou, Xiaoyan Gu, Bo Li, Weiping Wang
Abstract
Dataset distillation, also known as dataset condensation, offers a possibility for compressing a large-scale dataset into a small-scale one (i.e., distilled dataset) while achieving similar performance during model training. This method effectively tackles the challenges of training efficiency and storage cost posed by the large-scale dataset. Existing dataset distillation methods can be categorized into Optimization-Oriented (OO)-based and Distribution-Matching (DM)-based methods. Since OO-based methods require bi-level optimization to alternately optimize the model and the distilled data, they face challenges due to high computational overhead in practical applications. Thus, DM-based methods have emerged as an alternative by aligning the prototypes of the distilled data to those of the original data. Although efficient, these methods overlook the diversity of the distilled data, which will limit the performance of evaluation tasks. In this paper, we propose a novel Diversified Semantic Distribution Matching (DSDM) approach for dataset distillation. To accurately capture semantic features, we first pre-train models for dataset distillation. Subsequently, we estimate the distribution of each category by calculating its prototype and covariance matrix, where the covariance matrix indicates the direction of semantic feature transformations for each category. Then, in addition to the prototypes, the covariance matrices are also matched to obtain more diversity for the distilled data. However, since the distilled data are optimized by multiple pre-trained models, the training process will fluctuate severely. Therefore, we match the distilled data of the current pre-trained model with the historical integrated prototypes. Experimental results demonstrate that our DSDM achieves state-of-the-art results on both image and speech datasets. Code is available at https://github.com/Li-Hongcheng/DSDM.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 769b8b52-23dc-48d8-b3c3-ba30e98e21f6Cited by top-tier papers8
- Hyperbolic Dataset DistillationWenyuan Li, Guang Li, Keisuke Maeda, Takahiro Ogawa et al.NeurIPS 2025 · 17 citations
- Statistics Caching Test-Time Adaptation for Vision-Language ModelsZenghao Guan, Yucan Zhou, Wu Liu, Xiaoyan GuNeurIPS 2025 · 5 citations
- GeoDM: Geometry-aware Distribution Matching for Dataset DistillationXuhui Li, Zhengquan Luo, Zihui Cui, Kai Zhao et al.ICML 2026 · 2 citations
- Diversity-Enhanced Distribution Alignment for Dataset DistillationHongcheng Li, Yucan Zhou, Xiaoyan Gu, Bo Li et al.ICCV 2025 · 1 citation
- TGDD: Trajectory Guided Dataset Distillation with Balanced DistributionFengli Ran, Xiao Pu, Bo Liu, Xiuli Bi et al.AAAI 2026
Related papers
- Improved Distribution Matching for Dataset CondensationGanlong Zhao, Guanbin Li, Yipeng Qin, Yizhou YuCVPR 2023
- Dataset Distillation of 3D Point Clouds via Distribution MatchingJae-Young Yim, Dongwook Kim, Jae-Young SimNeurIPS 2025
- DREAM: Efficient Dataset Distillation by Representative MatchingYanqing Liu, Jianyang Gu, Kai Wang, Zheng Zhu et al.ICCV 2023 · 114 citations
- Multimodal Distribution Matching for Vision-Language Dataset DistillationJongoh Jeong, Hoyong Kwon, Minseok Kim, Kuk-Jin YoonCVPR 2026 · 3 citations
- M3D: Dataset Condensation by Minimizing Maximum Mean DiscrepancyHansong Zhang, Shikun Li, Pengju Wang, Dan Zeng et al.AAAI 2024 · 63 citations
