Rethinking Long-tailed Dataset Distillation: A Uni-Level Framework with Unbiased Recovery and Relabeling
Xiao Cui, Yulei Qin, Xinyue Li, Wengang Zhou, Hongsheng Li, Houqiang Li
摘要
Dataset distillation creates a small distilled set that enables efficient training by capturing key information from the full dataset. While existing dataset distillation methods perform well on balanced datasets, they struggle under long-tailed distributions, where imbalanced class frequencies induce biased model representations and corrupt statistical estimates such as Batch Normalization (BN) statistics. In this paper, we rethink long-tailed dataset distillation by revisiting the limitations of trajectory-based methods, and instead adopt the statistical alignment perspective to jointly mitigate model bias and restore fair supervision. To this end, we introduce three dedicated components that enable unbiased recovery of distilled images and soft relabeling: (1) enhancing expert models (an observer model for recovery and a teacher model for relabeling) to enable reliable statistics estimation and soft-label generation; (2) recalibrating BN statistics via a full forward pass with dynamically adjusted momentum to reduce representation skew; (3) initializing synthetic images by incrementally selecting high-confidence and diverse augmentations via a multi-round mechanism that promotes coverage and diversity. Extensive experiments on four long-tailed benchmarks show consistent improvements over state-of-the-art methods across varying degrees of class imbalance. Notably, our approach improves top-1 accuracy by 15.6% on CIFAR-100-LT and 11.8% on Tiny-ImageNet-LT under IPC=10 and IF=10.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Optimizing Distributional Geometry Alignment with Optimal Transport for Generative Dataset DistillationXiao Cui, Yulei Qin, Wengang Zhou, Hongsheng Li 等NeurIPS 2025 · 被引用 5 次
- Geometry-Aware Dataset Condensation for Diffusion Model TrainingXiao Cui, Yulei Qin, Mo Zhu, Wengang Zhou 等ICML 2026
它引用的顶会 Paper26
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 被引用 684 次
- Scaling Up Dataset Distillation to ImageNet-1K with Constant MemoryJustin Cui, Ruochen Wang, Si Si, Cho-Jui HsiehICML 2023 · 被引用 223 次
- Dataset Distillation by Matching Training TrajectoriesGeorge Cazenavette, Tongzhou Wang, Antonio Torralba, Alexei A. Efros 等CVPR 2022 · 被引用 198 次
- Squeeze, Recover and Relabel: Dataset Condensation at ImageNet Scale From A New PerspectiveZeyuan Yin, Eric P. Xing, Zhiqiang ShenNeurIPS 2023 · 被引用 180 次
- Towards Lossless Dataset Distillation via Difficulty-Aligned Trajectory MatchingZiyao Guo, Kai Wang, George Cazenavette, Hui Li 等ICLR 2024 · 被引用 142 次
相关 Paper
- Distilling Long-tailed DatasetsZhenghao Zhao, Haoxuan Wang, Yuzhang Shang, Kai Wang 等CVPR 2025
- Distilling Balanced Knowledge from a Biased TeacherSeonghak KimCVPR 2026 · 被引用 1 次
- Rectifying Soft-Label Entangled Bias in Long-Tailed Dataset DistillationChenyang Jiang, Hang Zhao, Xinyu Zhang, Zhengcen Li 等NeurIPS 2025 · 被引用 1 次
- Trust-calibrated Collaborative Learning for Long-Tailed Visual RecognitionHao Zhou, Tingjin LuoCVPR 2026
- Self Supervision to Distillation for Long-Tailed Visual RecognitionTianhao Li, Limin Wang, Gangshan WuICCV 2021 · 被引用 122 次
