Curriculum Coarse-to-Fine Selection for High-IPC Dataset Distillation
Yanda Chen, Gongwei Chen, Miao Zhang, Weili Guan, Liqiang Nie
摘要
Dataset distillation (DD) excels in synthesizing a small number of images per class (IPC) but struggles to maintain its effectiveness in high-IPC settings. Recent works on dataset distillation demonstrate that combining distilled and real data can mitigate the effectiveness decay. However, our analysis of the combination paradigm reveals that the current one-shot and independent selection mechanism induces an incompatibility issue between distilled and real images. To address this issue, we introduce a novel curriculum coarse-to-fine selection (CCFS) method for efficient high-IPC dataset distillation. CCFS employs a curriculum selection framework for real data selection, where we leverage a coarse-to-fine strategy to select appropriate real data based on the current synthetic dataset in each curriculum. Extensive experiments validate CCFS, surpassing the state-of-the-art by +6.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Beyond Random: Automatic Inner-loop Optimization in Dataset DistillationMuquan Li, Hang Gou, Dongyang Zhang, Shuang Liang 等NeurIPS 2025 · 被引用 8 次
- Balanced Dataset Distillation via Modeling Multiple Visual Pattern DistributionGuanghui Shi, Xuefeng Liang, Qixiang WenCVPR 2026 · 被引用 1 次
- Mind Your Margin and Boundary: Are Your Distilled Datasets Truly Robust?Muquan Li, Yingyi Ma, Yihong Huang, Hang Gou 等ICML 2026
- Correspondence Coverage Matters for Multi-Modal Dataset DistillationZhuohang Dang, Minnan Luo, Chengyou Jia, Hangwei Qian 等AAAI 2026
- Beyond Soft Label: Dataset Distillation via Orthogonal Gradient MatchingDeyu Bo, Xinchao WangCVPR 2026
它引用的顶会 Paper24
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 被引用 784 次
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 被引用 684 次
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 被引用 390 次
- Dataset Meta-Learning from Kernel Ridge-RegressionTimothy Nguyen, Zhourong Chen, Jaehoon LeeICLR 2021 · 被引用 307 次
- GLISTER: Generalization based Data Subset Selection for Efficient and Robust LearningKrishnaTeja Killamsetty, Durga Sivasubramanian, Ganesh Ramakrishnan, Rishabh K. IyerAAAI 2021 · 被引用 300 次
相关 Paper
- SelMatch: Effectively Scaling Up Dataset Distillation via Selection-Based Initialization and Partial Updates by Trajectory MatchingYongmin Lee, Hye Won ChungICML 2024 · 被引用 26 次
- DREAM: Efficient Dataset Distillation by Representative MatchingYanqing Liu, Jianyang Gu, Kai Wang, Zheng Zhu 等ICCV 2023 · 被引用 114 次
- Sequential Subset Matching for Dataset DistillationJiawei Du, Qin Shi, Joey Tianyi ZhouNeurIPS 2023 · 被引用 52 次
- Data Distillation Can Be Like Vodka: Distilling More Times For Better QualityXuxi Chen, Yu Yang, Zhangyang Wang, Baharan MirzasoleimanICLR 2024 · 被引用 19 次
- Large Scale Dataset Distillation with Domain ShiftNoel Loo, Alaa Maalouf, Ramin M. Hasani, Mathias Lechner 等ICML 2024 · 被引用 9 次
