Balanced Dataset Distillation via Modeling Multiple Visual Pattern Distribution
Guanghui Shi, Xuefeng Liang, Qixiang Wen
摘要
Dataset Distillation (DD) aims to compress large-scale datasets into a small number of condensed Images Per Class (IPC), enabling efficient network training. Previous coreset selection and synthetic-based DD methods achieve reasonable performance. However, our in-depth investigation reveals that existing methods share a common issue: pattern imbalance. Specifically, they either overemphasize class-general patterns representing the majority of each class or focus on fewer marginal patterns critical for model generalization. To address this issue, we propose a novel framework, Balanced Patterns Selection (BPS). Unlike prior methods that assume each class forms a single cluster, BPS models the multiple visual pattern distribution within each class via a hierarchical semantic structure inherent to the dataset. It then selects two complementary subsets in a balanced manner from the center (class-general patterns) and the margins (marginal patterns) of each pattern, producing a pattern-balanced coreset. Theoretically, we prove that the BPS-selected coreset aligns with the original dataset in both information coverage and performance. Moreover, its model-agnostic selection nature ensures cross-architecture generalization, while the Optimize-Once-for-All-IPCs property guarantees efficiency. Extensive experiments on four benchmarks demonstrate that BPS significantly outperforms existing state-ofthe-art methods. The source code is available at: https: //github.com/BeCarefulOfYournaoke/BPS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 被引用 684 次
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 被引用 494 次
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman 等ICLR 2020 · 被引用 462 次
- Coresets via Bilevel Optimization for Continual Learning and StreamingZalán Borsos, Mojmir Mutny, Andreas KrauseNeurIPS 2020 · 被引用 320 次
相关 Paper
- Diversified Semantic Distribution Matching for Dataset DistillationHongcheng Li, Yucan Zhou, Xiaoyan Gu, Bo Li 等ACM MM 2024 · 被引用 10 次
- Adaptive Dataset QuantizationMuquan Li, Dongyang Zhang, Qiang Dong, Xiurui Xie 等AAAI 2025 · 被引用 9 次
- Curriculum Coarse-to-Fine Selection for High-IPC Dataset DistillationYanda Chen, Gongwei Chen, Miao Zhang, Weili Guan 等CVPR 2025
- DREAM: Efficient Dataset Distillation by Representative MatchingYanqing Liu, Jianyang Gu, Kai Wang, Zheng Zhu 等ICCV 2023 · 被引用 114 次
- Data Distillation Can Be Like Vodka: Distilling More Times For Better QualityXuxi Chen, Yu Yang, Zhangyang Wang, Baharan MirzasoleimanICLR 2024 · 被引用 19 次
