On the Size and Approximation Error of Distilled Datasets
Alaa Maalouf, Murad Tukan, Noel Loo, Ramin M. Hasani, Mathias Lechner, Daniela Rus
Abstract
Dataset Distillation is the task of synthesizing small datasets from large ones while still retaining comparable predictive accuracy to the original uncompressed dataset. Despite significant empirical progress in recent years, there is little understanding of the theoretical limitations/guarantees of dataset distillation, specifically, what excess risk is achieved by distillation compared to the original dataset, and how large are distilled datasets? In this work, we take a theoretical view on kernel ridge regression (KRR) based methods of dataset distillation such as Kernel Inducing Points. By transforming ridge regression in random Fourier features (RFF) space, we provide the first proof of the existence of small (size) distilled datasets and their corresponding excess risk for shift-invariant kernels. We prove that a small set of instances exists in the original input space such that its solution in the RFF space coincides with the solution of the original data. We further show that a KRR solution can be generated using this distilled set of instances which gives an approximation towards the KRR solution optimized on the full input data. The size of this set is linear in the dimension of the RFF space of the input set or alternatively near linear in the number of effective degrees of freedom, which is a function of the kernel, number of datapoints, and the regularization parameter λ. The error bound of this distilled set is also a function of λ. We verify our bounds analytically and empirically.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0b11f235-eefb-4b27-8683-310ae2d2251aCited by top-tier papers3
- Large Scale Dataset Distillation with Domain ShiftNoel Loo, Alaa Maalouf, Ramin M. Hasani, Mathias Lechner et al.ICML 2024 · 9 citations
- Dataset Distillation Efficiently Encodes Low-Dimensional Representations from Gradient-Based Learning of Non-Linear TasksYuri Kinoshita, Naoki Nishikawa, Taro ToyoizumiICML 2026
- Monotonic Variational Gaussian Process for Efficient Data CollectionDonghyun Lee, Young Myoung KoICML 2026
Builds on16
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 390 citations
- Coresets via Bilevel Optimization for Continual Learning and StreamingZalán Borsos, Mojmir Mutny, Andreas KrauseNeurIPS 2020 · 320 citations
- Dataset Distillation with Infinitely Wide Convolutional NetworksTimothy Nguyen, Roman Novak, Lechao Xiao, Jaehoon LeeNeurIPS 2021 · 313 citations
Related papers
- Provable and Efficient Dataset Distillation for Kernel Ridge RegressionYilan Chen, Wei Huang, Lily WengNeurIPS 2024 · 8 citations
- Efficient Dataset Distillation using Random Feature ApproximationNoel Loo, Ramin M. Hasani, Alexander Amini, Daniela RusNeurIPS 2022 · 156 citations
- Dataset Meta-Learning from Kernel Ridge-RegressionTimothy Nguyen, Zhourong Chen, Jaehoon LeeICLR 2021 · 307 citations
- Implicit Regularization of Random Feature ModelsArthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler et al.ICML 2020 · 83 citations
- Random Fourier Features via Fast Surrogate Leverage Weighted SamplingFanghui Liu, Xiaolin Huang, Yudong Chen, Jie Yang et al.AAAI 2020 · 21 citations
