Hard Labels In! Rethinking the Role of Hard Labels in Mitigating Local Semantic Drift
Jiacheng Cui, Bingkui Tong, Xinyue Bi, Xiaohan Zhao, Jiacheng Liu, Zhiqiang Shen
Abstract
Soft labels from teacher models are a de facto practice for knowledge transfer and large-scale dataset distillation (e.g., SReL, LPLD). However, when we limit the number of crops per image to reduce the substantial cost of storing precomputed soft labels, these methods suffer severely from local semantic drift: visually ambiguous crops can cause soft supervision to deviate from the image-level ground-truth semantics, leading to persistent errors and a train–test distribution mismatch. We revisit the overlooked role of hard labels and show that, when properly integrated, they can act as a content-invariant semantic anchor that calibrates such drift. We theoretically analyze the emergence of drift under sparse soft-label supervision and demonstrate that hybridizing hard and soft labels restores alignment between visual content and semantic supervision. Building on this insight, we propose a new training paradigm, Hard Label for Alleviating Local Semantic Drift (HALD), which uses hard labels as intermediate corrective signals while preserving the fine-grained benefits of soft labels. Extensive experiments on dataset distillation and large-scale classification benchmarks show consistent generalization improvements. On ImageNet-1K, our method achieves 42.7% accuracy with only 285M soft-label storage (reduces by 100), outperforming prior state-of-the-art LPLD by 9.0%. Code is available at https://github.com/Jiacheng8/HALD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 531d8995-d5e3-4198-be23-e2fa151006a4Builds on34
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 390 citations
- Dataset Distillation with Infinitely Wide Convolutional NetworksTimothy Nguyen, Roman Novak, Lechao Xiao, Jaehoon LeeNeurIPS 2021 · 313 citations
- Dataset Distillation using Neural Feature RegressionYongchao Zhou, Ehsan Nezhadarya, Jimmy BaNeurIPS 2022 · 234 citations
- Dataset Condensation via Efficient Synthetic-Data ParameterizationJang-Hyun Kim, Jinuk Kim, Seong Joon Oh, Sangdoo Yun et al.ICML 2022 · 234 citations
Related papers
- FerKD: Surgical Label Adaptation for Efficient DistillationZhiqiang ShenICCV 2023 · 7 citations
- Rethinking Dataset Distillation: Hard Truths about Soft LabelsPriyam Dey, Aditya Sahdev, Sunny Bhati, Konda Reddy Mopuri et al.CVPR 2026 · 2 citations
- Scale-Equivalent Distillation for Semi-Supervised Object DetectionQiushan Guo, Yao Mu, Jianyu Chen, Tianqi Wang et al.CVPR 2022 · 40 citations
- Self Supervision to Distillation for Long-Tailed Visual RecognitionTianhao Li, Limin Wang, Gangshan WuICCV 2021 · 122 citations
- A Label is Worth A Thousand Images in Dataset DistillationTian Qin, Zhiwei Deng, David Alvarez-MelisNeurIPS 2024 · 39 citations
