Is Bin Generation Indispensable? A Bin-Generation-Free Dataset Quantization via Semantic Perspective
Maijie Deng, Yuhua Li, Yixiong Zou, Yao Wu, Chenru Ma
摘要
Dataset quantization has recently emerged as a promising solution for mitigating the computational and memory challenges of large-scale datasets. However, existing approaches rely on a bin generation step that is computationally expensive and inefficient for large-scale datasets. Moreover, a fixed drop ratio in its patch dropping step fails to adapt to the diverse redundancy levels across samples, which degrades the representational quality of the quantized coreset. To address these limitations, we present Bin-Generation-Free Dataset Quantization (BGFDQ), a fully restructured framework that incorporates a simple yet effective KNN-based neighbor identification and neighboraware coreset selection strategy. We theoretically demonstrate that the proposed selection strategy achieves superior sampling efficiency compared to bin-generation-based methods. Additionally, we introduce an adaptive patch dropping strategy to further enhance the quality of the quantized dataset. Extensive experiments on four image classification benchmarks show that BGFDQ consistently outperforms state-of-the-art baselines. In particular, we achieve up to 5% validation accuracy improvement on CIFAR-100. Moreover, our framework successfully scales to datasets containing up to 10 5 same-class samples while existing bin-generation-based approaches fail due to memory constraints. Code is available at https://github.com/MaijieDeng/BGFDQ.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang 等NeurIPS 2025 · 被引用 949 次
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 被引用 806 次
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 被引用 684 次
- GRAD-MATCH: Gradient Matching based Data Subset Selection for Efficient Deep Model TrainingKrishnaTeja Killamsetty, Durga Sivasubramanian, Ganesh Ramakrishnan, Abir De 等ICML 2021 · 被引用 305 次
- Dataset Condensation via Efficient Synthetic-Data ParameterizationJang-Hyun Kim, Jinuk Kim, Seong Joon Oh, Sangdoo Yun 等ICML 2022 · 被引用 234 次
相关 Paper
- Dataset Color Quantization: A Training-Oriented Framework for Dataset-Level CompressionYU CHENYUE, Lingao Xiao, Jinhong Deng, Ivor Tsang 等ICLR 2026
- Adaptive Dataset QuantizationMuquan Li, Dongyang Zhang, Qiang Dong, Xiurui Xie 等AAAI 2025 · 被引用 9 次
- Post Training Quantization for Efficient Dataset CondensationLinh-Tam Tran, Sung-Ho BaeAAAI 2026
- Dataset QuantizationDaquan Zhou, Kai Wang, Jianyang Gu, Xiangyu Peng 等ICCV 2023 · 被引用 65 次
- Unleashing the Full Potential of Product Quantization for Large-Scale Image RetrievalYu Liang, Shiliang Zhang, Li Ken Li, Xiaoyu WangNeurIPS 2023 · 被引用 5 次
