Embarrassingly Simple Dataset Distillation
Yunzhen Feng, Shanmukha Ramakrishna Vedantam, Julia Kempe
Abstract
Training of large-scale models in general requires enormous amounts of traning data. Dataset distillation aims to extract a small set of synthetic training samples from a large dataset with the goal of achieving competitive performance on test data when trained on this sample, thus reducing both dataset size and training time. In this work, we tackle dataset distillation at its core by treating it directly as a bilevel optimization problem. Re-examining the foundational back-propagation through time method, we study the pronounced variance in the gradients, computational burden, and long-term dependencies. We introduce an improved method: Random Truncated Backpropagation Through Time (RaT-BPTT) to address them. RaT-BPTT incorporates a truncation coupled with a random window, effectively stabilizing the gradients and speeding up the optimization while covering long dependencies. This allows us to establish new dataset distillation state-of-the-art for a variety of standard dataset benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f1dc364-0acf-46aa-97e2-24abc70aa41cCited by top-tier papers12
- Diversity-Driven Synthesis: Enhancing Dataset Distillation through Directed Weight AdjustmentJiawei Du, Xin Zhang, Juncheng Hu, Wenxin Huang et al.NeurIPS 2024 · 43 citations
- TD3: Tucker Decomposition Based Dataset Distillation Method for Sequential RecommendationJiaqing Zhang, Mingjia Yin, Hao Wang, Yawen Li et al.WWW 2025 · 17 citations
- Teaching Models to Teach Themselves: Reasoning at the Edge of LearnabilityShobhita Sundaram, John Quan, Ariel Kwiatkowski, Kartik Ahuja et al.ICML 2026 · 16 citations
- Beyond Random: Automatic Inner-loop Optimization in Dataset DistillationMuquan Li, Hang Gou, Dongyang Zhang, Shuang Liang et al.NeurIPS 2025 · 8 citations
- Provable and Efficient Dataset Distillation for Kernel Ridge RegressionYilan Chen, Wei Huang, Lily WengNeurIPS 2024 · 8 citations
Builds on26
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 390 citations
- Dataset Distillation with Infinitely Wide Convolutional NetworksTimothy Nguyen, Roman Novak, Lechao Xiao, Jaehoon LeeNeurIPS 2021 · 313 citations
- Dataset Meta-Learning from Kernel Ridge-RegressionTimothy Nguyen, Zhourong Chen, Jaehoon LeeICLR 2021 · 307 citations
Related papers
- Dataset Distillation using Neural Feature RegressionYongchao Zhou, Ehsan Nezhadarya, Jimmy BaNeurIPS 2022 · 234 citations
- Data Distillation Can Be Like Vodka: Distilling More Times For Better QualityXuxi Chen, Yu Yang, Zhangyang Wang, Baharan MirzasoleimanICLR 2024 · 19 citations
- Rectified Decoupled Dataset Distillation: A Closer Look for Fair and Comprehensive EvaluationXinhao Zhong, Shuoyang Sun, Xulin Gu, Chenyang Zhu et al.ICLR 2026 · 2 citations
- An Adaptive Sampling Framework for Diffusion-based Dataset Distillation with High Fidelity and DiversitySunbeom Jeong, Sehwan Kim, Hyeonggeun Han, Hyungjun Joo et al.AAAI 2026
- Towards Stable and Storage-efficient Dataset Distillation: Matching Convexified TrajectoryWenliang Zhong, Haoyu Tang, Qinghai Zheng, Mingzhu Xu et al.CVPR 2025
