Asymmetric Synthetic Data Update for Domain Incremental Dataset Distillation
Minyoung Oh, Jae-Young Sim
Abstract
Dataset distillation (DD) attempts to construct a compact synthetic dataset that serves as a proxy for a large real dataset under a fixed storage budget, thereby reducing the storage burden and training costs. Prior works assume the full dataset is available upfront which is distilled at once, although real datasets are collected incrementally over time in practice. To alleviate this gap, we introduce a new problem setting, Domain Incremental Dataset Distillation, that continually distills datasets from different domains into a single synthetic dataset. The conventional DD sequentially processes arriving datasets in order, overwriting the old knowledge with new one, causing catastrophic forgetting problem. To overcome this drawback, we propose Asymmetric Synthetic Data Update strategy that adjusts the per-sample update rates for synthetic dataset while balancing the stability-plasticity trade-off. Specifically, we design a bi-level optimization method based on meta-learning framework to estimate the optimal update rates, which allows each sample to focus on either stability or plasticity, thereby striking a balance between them. Experimental results demonstrate that our approach effectively mitigates the catastrophic forgetting and achieves superior performance of DD across continually incoming datasets compared with existing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd8349ed-e1fd-4987-88e4-973c23773e7cBuilds on14
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 390 citations
- Dataset Condensation via Efficient Synthetic-Data ParameterizationJang-Hyun Kim, Jinuk Kim, Seong Joon Oh, Sangdoo Yun et al.ICML 2022 · 234 citations
- Dataset Distillation by Matching Training TrajectoriesGeorge Cazenavette, Tongzhou Wang, Antonio Torralba, Alexei A. Efros et al.CVPR 2022 · 198 citations
- Generative Teaching Networks: Accelerating Neural Architecture Search by Learning to Generate Synthetic Training DataFelipe Petroski Such, Aditya Rawal, Joel Lehman, Kenneth O. Stanley et al.ICML 2020 · 180 citations
Related papers
- Continual Learning with Semi-supervised Contrastive Distillation for Incremental Neural Machine TranslationYunlong Liang, Fandong Meng, Jiaan Wang, Jinan Xu et al.ACL 2024 · 7 citations
- MGDD: A Meta Generator for Fast Dataset DistillationSonghua Liu, Xinchao WangNeurIPS 2023 · 13 citations
- Incremental Learning in Online ScenarioJiangpeng He, Runyu Mao, Zeman Shao, Fengqing ZhuCVPR 2020
- DFIL: Deepfake Incremental Learning by Exploiting Domain-invariant Forgery CluesKun Pan, Yifang Yin, Yao Wei, Feng Lin et al.ACM MM 2023 · 35 citations
- Always Be Dreaming: A New Approach for Data-Free Class-Incremental LearningJames Seale Smith, Yen-Chang Hsu, Jonathan C. Balloch, Yilin Shen et al.ICCV 2021 · 208 citations
