TGDD: Trajectory Guided Dataset Distillation with Balanced Distribution
Fengli Ran, Xiao Pu, Bo Liu, Xiuli Bi, Bin Xiao
Abstract
Dataset distillation compresses large datasets into compact synthetic ones to reduce storage and computational costs. Among various approaches, distribution matching (DM)-based methods have attracted attention for their high efficiency. However, they often overlook the evolution of feature representations during training, which limits the expressiveness of synthetic data and weakens downstream performance. To address this issue, we propose Trajectory Guided Dataset Distillation (TGDD), which reformulates distribution matching as a dynamic alignment process along the model’s training trajectory. At each training stage, TGDD captures evolving semantics by aligning the feature distribution between the synthetic and original dataset. Meanwhile, it introduces a distribution constraint regularization to reduce class overlap. This design helps synthetic data preserve both semantic diversity and representativeness, improving performance in downstream tasks. Without additional optimization overhead, TGDD achieves a favorable balance between performance and efficiency. Experiments on ten datasets demonstrate that TGDD achieves state-of-the-art performance, notably a 5.0% accuracy gain on high-resolution benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 41e64f83-de2b-47f6-9dd6-30892e8aacbaBuilds on17
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 390 citations
- Scaling Up Dataset Distillation to ImageNet-1K with Constant MemoryJustin Cui, Ruochen Wang, Si Si, Cho-Jui HsiehICML 2023 · 223 citations
- Dataset Distillation by Matching Training TrajectoriesGeorge Cazenavette, Tongzhou Wang, Antonio Torralba, Alexei A. Efros et al.CVPR 2022 · 198 citations
- Generative Teaching Networks: Accelerating Neural Architecture Search by Learning to Generate Synthetic Training DataFelipe Petroski Such, Aditya Rawal, Joel Lehman, Kenneth O. Stanley et al.ICML 2020 · 180 citations
- Privacy for Free: How does Dataset Condensation Help Privacy?Tian Dong, Bo Zhao, Lingjuan LyuICML 2022 · 154 citations
Related papers
- DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion ModelsQichao Wang, Yunhong Lu, Hengyuan Cao, Junyi Zhang et al.CVPR 2026 · 4 citations
- Diversified Semantic Distribution Matching for Dataset DistillationHongcheng Li, Yucan Zhou, Xiaoyan Gu, Bo Li et al.ACM MM 2024 · 10 citations
- DREAM: Efficient Dataset Distillation by Representative MatchingYanqing Liu, Jianyang Gu, Kai Wang, Zheng Zhu et al.ICCV 2023 · 114 citations
- Towards Lossless Dataset Distillation via Difficulty-Aligned Trajectory MatchingZiyao Guo, Kai Wang, George Cazenavette, Hui Li et al.ICLR 2024 · 142 citations
- Minimizing the Accumulated Trajectory Error to Improve Dataset DistillationJiawei Du, Yidi Jiang, Vincent Y. F. Tan, Joey Tianyi Zhou et al.CVPR 2023
