Enhancing Dataset Distillation via Non-Critical Region Refinement
Minh-Tuan Tran, Trung Le, Xuan-May Le, Thanh-Toan Do, Dinh Q. Phung
Abstract
Dataset distillation has become a popular method for compressing large datasets into smaller, more efficient representations while preserving critical information for model training. Data features are broadly categorized into two types: instance-specific features, which capture unique, fine-grained details of individual examples, and classgeneral features, which represent shared, broad patterns across a class. However, previous approaches often struggle to balance these features-some focus solely on classgeneral patterns, neglecting finer instance details, while others prioritize instance-specific features, overlooking the shared characteristics essential for class-level understanding. In this paper, we introduce the Non-Critical Region Refinement Dataset Distillation (NRR-DD) method, which preserves instance-specific details and fine-grained regions in synthetic data while enriching non-critical regions with class-general information. This approach enables models to leverage all pixel information, capturing both feature types and enhancing overall performance. Additionally, we present Distance-Based Representative (DBR) knowledge transfer, which eliminates the need for soft labels in training by relying on the distance between synthetic data predictions and one-hot encoded labels. Experimental results show that NRR-DD achieves state-of-the-art performance on both small-and large-scale datasets. Furthermore, by storing only two distances per instance, our method delivers comparable results across various settings. The code is available at https://github.com/tmtuan1307/ NRR-DD .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8e6c73af-e31e-425d-ab7b-33d96983d164Cited by top-tier papers9
- Fixed Anchors Are Not Enough: Dynamic Retrieval and Persistent Homology for Dataset DistillationMuquan Li, Hang Gou, Yingyi Ma, Rongzheng Wang et al.CVPR 2026 · 11 citations
- Beyond Random: Automatic Inner-loop Optimization in Dataset DistillationMuquan Li, Hang Gou, Dongyang Zhang, Shuang Liang et al.NeurIPS 2025 · 8 citations
- Efficient Multimodal Dataset Distillation via Generative ModelsZhenghao Zhao, Haoxuan Wang, Junyi Wu, Yuzhang Shang et al.NeurIPS 2025 · 7 citations
- Balanced Dataset Distillation via Modeling Multiple Visual Pattern DistributionGuanghui Shi, Xuefeng Liang, Qixiang WenCVPR 2026 · 1 citation
- Multimodal Dataset Distillation Made Simple by Prototype-Guided Data SynthesisJunhyeok Choi, Sangwoo Mo, Minwoo ChaeICLR 2026
Builds on15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- Scaling Up Dataset Distillation to ImageNet-1K with Constant MemoryJustin Cui, Ruochen Wang, Si Si, Cho-Jui HsiehICML 2023 · 223 citations
- Dataset Distillation by Matching Training TrajectoriesGeorge Cazenavette, Tongzhou Wang, Antonio Torralba, Alexei A. Efros et al.CVPR 2022 · 198 citations
Related papers
- Optimizing Distributional Geometry Alignment with Optimal Transport for Generative Dataset DistillationXiao Cui, Yulei Qin, Wengang Zhou, Hongsheng Li et al.NeurIPS 2025 · 5 citations
- TGDD: Trajectory Guided Dataset Distillation with Balanced DistributionFengli Ran, Xiao Pu, Bo Liu, Xiuli Bi et al.AAAI 2026
- A Label is Worth A Thousand Images in Dataset DistillationTian Qin, Zhiwei Deng, David Alvarez-MelisNeurIPS 2024 · 39 citations
- Going Beyond Feature Similarity: Effective Dataset distillation based on Class-aware Conditional Mutual InformationXinhao Zhong, Bin Chen, Hao Fang, Xulin Gu et al.ICLR 2025
- Breaking Class Barriers: Efficient Dataset Distillation via Inter-Class Feature CompensatorXin Zhang, Jiawei Du, Ping Liu, Joey Tianyi ZhouICLR 2025
