Robust Dataset Condensation using Supervised Contrastive Learning
Nicole Hee-Yeon Kim, Hwanjun Song
Abstract
Dataset condensation aims to compress large dataset into smaller synthetic set while preserving the essential representations needed for effective model training. However, existing methods show severe performance degradation when applied to noisy datasets. To address this, we present robust dataset condensation (RDC), an end-to-end method that mitigates noise to generate a clean and robust synthetic set, without requiring separate noise-reduction preprocessing steps. RDC refines the condensation process by integrating contrastive learning tailored for robust condensation, named golden MixUp contrast. It uses synthetic samples to sharpen class boundaries and to mitigate noisy representations, while its augmentation strategy compensates for the limited size of the synthetic set by identifying clean samples from noisy training data, enriching synthetic images with real-data diversity. We evaluate RDC against existing condensation methods and a conventional approach that first applies noise cleaning algorithms to the dataset before performing condensation. Extensive experiments show that RDC outperforms other approaches on CIFAR-10/100 across different types of noise, including asymmetric, symmetric, and real-world noise. Code is available at https: //github.com/DISL-Lab/RDC-ICCV2025.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0c613a73-4f0c-4ddf-b9d1-c13ebec73666Builds on25
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 1,326 citations
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
Related papers
- Dataset Condensation with Contrastive SignalsSaehyung Lee, Sanghyuk Chun, Sangwon Jung, Sangdoo Yun et al.ICML 2022 · 139 citations
- Learning from Noisy Data with Robust Representation LearningJunnan Li, Caiming Xiong, Steven C. H. HoiICCV 2021 · 140 citations
- Noise-Optimized Distribution Distillation for Dataset CondensationTongfei Liu, Yufan Liu, Bing Li, Weiming Hu et al.ACM MM 2025
- M3D: Dataset Condensation by Minimizing Maximum Mean DiscrepancyHansong Zhang, Shikun Li, Pengju Wang, Dan Zeng et al.AAAI 2024 · 63 citations
- An Efficient Dataset Condensation Plugin and Its Application to Continual LearningEnneng Yang, Li Shen, Zhenyi Wang, Tongliang Liu et al.NeurIPS 2023 · 49 citations
