Dataset Distillation via Factorization
Songhua Liu, Kai Wang, Xingyi Yang, Jingwen Ye, Xinchao Wang
Abstract
In this paper, we study dataset distillation (DD), from a novel perspective and introduce a dataset factorization approach, termed HaBa, which is a plug-and-play strategy portable to any existing DD baseline. Unlike conventional DD approaches that aim to produce distilled and representative samples, HaBa explores decomposing a dataset into two components: data Hallucination networks and Bases, where the latter is fed into the former to reconstruct image samples. The flexible combinations between bases and hallucination networks, therefore, equip the distilled data with exponential informativeness gain, which largely increase the representation capability of distilled datasets. To furthermore increase the data efficiency of compression results, we further introduce a pair of adversarial contrastive constraints on the resultant hallucination networks and bases, which increase the diversity of generated images and inject more discriminant information into the factorization. Extensive comparisons and experiments demonstrate that our method can yield significant improvement on downstream classification tasks compared with previous state of the arts, while reducing the total number of compressed parameters by up to 65%. Moreover, distilled datasets by our approach also achieve 10% higher accuracy than baseline methods in cross-architecture generalization. Our code is available here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers81
- One-for-All: Bridge the Gap Between Heterogeneous Architectures in Knowledge DistillationZhiwei Hao, Jianyuan Guo, Kai Han, Yehui Tang et al.NeurIPS 2023 · 205 citations
- Squeeze, Recover and Relabel: Dataset Condensation at ImageNet Scale From A New PerspectiveZeyuan Yin, Eric P. Xing, Zhiqiang ShenNeurIPS 2023 · 180 citations
- Towards Lossless Dataset Distillation via Difficulty-Aligned Trajectory MatchingZiyao Guo, Kai Wang, George Cazenavette, Hui Li et al.ICLR 2024 · 142 citations
- DREAM: Efficient Dataset Distillation by Representative MatchingYanqing Liu, Jianyang Gu, Kai Wang, Zheng Zhu et al.ICCV 2023 · 114 citations
- Diffusion Model as Representation LearnerXingyi Yang, Xinchao WangICCV 2023 · 100 citations
Builds on17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 390 citations
Related papers
- Boost Self-Supervised Dataset Distillation via Parameterization, Predefined Augmentation, and ApproximationSheng-Feng Yu, Jia-Jiun Yao, Wei-Chen ChiuICLR 2025
- DIVER: Diving Deeper into Distilled Data via Expressive Semantic RecoveryQianxin Xia, Zhiyong Shu, Wenbo Jiang, Jiawei Du et al.ICML 2026
- An Adaptive Sampling Framework for Diffusion-based Dataset Distillation with High Fidelity and DiversitySunbeom Jeong, Sehwan Kim, Hyeonggeun Han, Hyungjun Joo et al.AAAI 2026
- Adaptive Dataset QuantizationMuquan Li, Dongyang Zhang, Qiang Dong, Xiurui Xie et al.AAAI 2025 · 9 citations
- D4M: Dataset Distillation via Disentangled Diffusion ModelDuo Su, Junjie Hou, Weizhi Gao, Yingjie Tian et al.CVPR 2024 · 11 citations
