An Efficient Dataset Condensation Plugin and Its Application to Continual Learning
Enneng Yang, Li Shen, Zhenyi Wang, Tongliang Liu, Guibing Guo
Abstract
Dataset condensation (DC) distills a large real-world dataset into a small synthetic dataset, with the goal of training a network from scratch on the latter that performs similarly to the former. State-of-the-art (SOTA) DC methods have achieved satisfactory results through techniques such as accuracy, gradient, training trajectory, or distribution matching. However, these works all perform matching in the high-dimension pixel space, ignoring that natural images are usually locally connected and have lower intrinsic dimensions, resulting in low condensation efficiency. In this work, we propose a simple-yet-efficient dataset condensation plugin that matches the raw and synthetic datasets in a low-dimensional manifold. Specifically, our plugin condenses raw images into two low-rank matrices instead of parameterized image matrices. Our plugin can be easily incorporated into existing DC methods, thereby containing richer raw dataset information at limited storage costs to improve the downstream applications’ performance. We verify on multiple public datasets that when the proposed plugin is combined with SOTA DC methods, the performance of the network trained on synthetic data is significantly improved compared to traditional DC methods. Moreover, when applying the DC methods as a plugin to continual learning tasks, we observed that our approach effectively mitigates catastrophic forgetting of old tasks under limited memory buffer constraints and avoids the problem of raw data privacy leakage.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- A Unified and General Framework for Continual LearningZhenyi Wang, Yan Li, Li Shen, Heng HuangICLR 2024 · 42 citations
- Hyperbolic Dataset DistillationWenyuan Li, Guang Li, Keisuke Maeda, Takahiro Ogawa et al.NeurIPS 2025 · 17 citations
- TD3: Tucker Decomposition Based Dataset Distillation Method for Sequential RecommendationJiaqing Zhang, Mingjia Yin, Hao Wang, Yawen Li et al.WWW 2025 · 17 citations
- Fixed Anchors Are Not Enough: Dynamic Retrieval and Persistent Homology for Dataset DistillationMuquan Li, Hang Gou, Yingyi Ma, Rongzheng Wang et al.CVPR 2026 · 11 citations
- Provable and Efficient Dataset Distillation for Kernel Ridge RegressionYilan Chen, Wei Huang, Lily WengNeurIPS 2024 · 8 citations
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
Related papers
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- Dataset Condensation with Contrastive SignalsSaehyung Lee, Sanghyuk Chun, Sangwon Jung, Sangdoo Yun et al.ICML 2022 · 139 citations
- Remember the Past: Distilling Datasets into Addressable Memories for Neural NetworksZhiwei Deng, Olga RussakovskyNeurIPS 2022 · 140 citations
- DataDAM: Efficient Dataset Distillation with Attention MatchingAhmad Sajedi, Samir Khaki, Ehsan Amjadian, Lucy Z. Liu et al.ICCV 2023 · 106 citations
- Slimmable Dataset CondensationSonghua Liu, Jingwen Ye, Runpeng Yu, Xinchao WangCVPR 2023
