Dataset Condensation with Contrastive Signals
Saehyung Lee, Sanghyuk Chun, Sangwon Jung, Sangdoo Yun, Sungroh Yoon
Abstract
Recent studies have demonstrated that gradient matching-based dataset synthesis, or dataset condensation (DC), methods can achieve state-of-the-art performance when applied to data-efficient learning tasks. However, in this study, we prove that the existing DC methods can perform worse than the random selection method when task-irrelevant information forms a significant part of the training dataset. We attribute this to the lack of participation of the contrastive signals between the classes resulting from the class-wise gradient matching strategy. To address this problem, we propose Dataset Condensation with Contrastive signals (DCC) by modifying the loss function to enable the DC methods to effectively capture the differences between classes. In addition, we analyze the new loss function in terms of training dynamics by tracking the kernel velocity. Furthermore, we introduce a bi-level warm-up strategy to stabilize the optimization. Our experimental results indicate that while the existing methods are ineffective for fine-grained image classification tasks, the proposed method can successfully generate informative synthetic datasets for the same tasks. Moreover, we demonstrate that the proposed method outperforms the baselines even on benchmark datasets such as SVHN, CIFAR-10, and CIFAR-100. Finally, we demonstrate the high applicability of the proposed method by applying it to continual learning tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers72
- Dataset Distillation using Neural Feature RegressionYongchao Zhou, Ehsan Nezhadarya, Jimmy BaNeurIPS 2022 · 234 citations
- Scaling Up Dataset Distillation to ImageNet-1K with Constant MemoryJustin Cui, Ruochen Wang, Si Si, Cho-Jui HsiehICML 2023 · 223 citations
- Dataset Distillation via FactorizationSonghua Liu, Kai Wang, Xingyi Yang, Jingwen Ye et al.NeurIPS 2022 · 190 citations
- Squeeze, Recover and Relabel: Dataset Condensation at ImageNet Scale From A New PerspectiveZeyuan Yin, Eric P. Xing, Zhiqiang ShenNeurIPS 2023 · 180 citations
- DataDAM: Efficient Dataset Distillation with Attention MatchingAhmad Sajedi, Samir Khaki, Ehsan Amjadian, Lucy Z. Liu et al.ICCV 2023 · 106 citations
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
Related papers
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- An Efficient Dataset Condensation Plugin and Its Application to Continual LearningEnneng Yang, Li Shen, Zhenyi Wang, Tongliang Liu et al.NeurIPS 2023 · 49 citations
- Sequential Subset Matching for Dataset DistillationJiawei Du, Qin Shi, Joey Tianyi ZhouNeurIPS 2023 · 52 citations
- Condensing Graphs via One-Step Gradient MatchingWei Jin, Xianfeng Tang, Haoming Jiang, Zheng Li et al.KDD 2022 · 68 citations
- CAFE: Learning to Condense Dataset by Aligning FeaturesKai Wang, Bo Zhao, Xiangyu Peng, Zheng Zhu et al.CVPR 2022 · 140 citations
