Distilling Long-tailed Datasets
Zhenghao Zhao, Haoxuan Wang, Yuzhang Shang, Kai Wang, Yan Yan
摘要
Dataset distillation aims to synthesize a small, informationrich dataset from a large one for efficient model training. However, existing dataset distillation methods struggle with long-tailed datasets, which are prevalent in real-world scenarios. By investigating the reasons behind this unexpected result, we identified two main causes: 1) The distillation process on imbalanced datasets develops biased gradients, leading to the synthesis of similarly imbalanced distilled datasets. 2) The experts trained on such datasets perform suboptimally on tail classes, resulting in misguided distillation supervision and poor-quality soft-label initialization. To address these issues, we first propose Distributionagnostic Matching to avoid directly matching the biased expert trajectories. It reduces the distance between the student and the biased expert trajectories and prevents the tail class bias from being distilled to the synthetic dataset. Moreover, we improve the distillation guidance with Expert Decoupling, which jointly matches the decoupled backbone and classifier to improve the tail class performance and initialize reliable soft labels. This work pioneers the field of longtailed dataset distillation, marking the first effective effort to distill long-tailed datasets. Our code will be made public at https://github.com/ichbill/LTDD .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Efficient Multimodal Dataset Distillation via Generative ModelsZhenghao Zhao, Haoxuan Wang, Junyi Wu, Yuzhang Shang 等NeurIPS 2025 · 被引用 7 次
- FairDD: Fair Dataset DistillationQihang Zhou, Shenhao Fang, Shibo He, Wenchao Meng 等NeurIPS 2025 · 被引用 3 次
- CaO2: Rectifying Inconsistencies in Diffusion-Based Dataset DistillationHaoxuan Wang, Zhenghao Zhao, Junyi Wu, Yuzhang Shang 等ICCV 2025 · 被引用 1 次
- Rectifying Soft-Label Entangled Bias in Long-Tailed Dataset DistillationChenyang Jiang, Hang Zhao, Xinyu Zhang, Zhengcen Li 等NeurIPS 2025 · 被引用 1 次
- Rethinking Long-tailed Dataset Distillation: A Uni-Level Framework with Unbiased Recovery and RelabelingXiao Cui, Yulei Qin, Xinyue Li, Wengang Zhou 等AAAI 2026 · 被引用 1 次
它引用的顶会 Paper23
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan 等ICLR 2020 · 被引用 1,496 次
- Balanced Meta-Softmax for Long-Tailed Visual RecognitionJiawei Ren, Cunjun Yu, Shunan Sheng, Xiao Ma 等NeurIPS 2020 · 被引用 861 次
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 被引用 684 次
- Dataset Meta-Learning from Kernel Ridge-RegressionTimothy Nguyen, Zhourong Chen, Jaehoon LeeICLR 2021 · 被引用 307 次
- Dataset Distillation using Neural Feature RegressionYongchao Zhou, Ehsan Nezhadarya, Jimmy BaNeurIPS 2022 · 被引用 234 次
相关 Paper
- Distilling Balanced Knowledge from a Biased TeacherSeonghak KimCVPR 2026 · 被引用 1 次
- Self Supervision to Distillation for Long-Tailed Visual RecognitionTianhao Li, Limin Wang, Gangshan WuICCV 2021 · 被引用 122 次
- Decoupled Contrastive Learning for Long-Tailed RecognitionShiyu Xuan, Shiliang ZhangAAAI 2024 · 被引用 29 次
- Distilling Virtual Examples for Long-tailed RecognitionYin-Yin He, Jianxin Wu, Xiu-Shen WeiICCV 2021 · 被引用 129 次
- Self-Supervised Aggregation of Diverse Experts for Test-Agnostic Long-Tailed RecognitionYifan Zhang, Bryan Hooi, Lanqing Hong, Jiashi FengNeurIPS 2022 · 被引用 214 次
