Is Bin Generation Indispensable? A Bin-Generation-Free Dataset Quantization via Semantic Perspective
Maijie Deng, Yuhua Li, Yixiong Zou, Yao Wu, Chenru Ma
Abstract
Dataset quantization has recently emerged as a promising solution for mitigating the computational and memory challenges of large-scale datasets. However, existing approaches rely on a bin generation step that is computationally expensive and inefficient for large-scale datasets. Moreover, a fixed drop ratio in its patch dropping step fails to adapt to the diverse redundancy levels across samples, which degrades the representational quality of the quantized coreset. To address these limitations, we present Bin-Generation-Free Dataset Quantization (BGFDQ), a fully restructured framework that incorporates a simple yet effective KNN-based neighbor identification and neighboraware coreset selection strategy. We theoretically demonstrate that the proposed selection strategy achieves superior sampling efficiency compared to bin-generation-based methods. Additionally, we introduce an adaptive patch dropping strategy to further enhance the quality of the quantized dataset. Extensive experiments on four image classification benchmarks show that BGFDQ consistently outperforms state-of-the-art baselines. In particular, we achieve up to 5% validation accuracy improvement on CIFAR-100. Moreover, our framework successfully scales to datasets containing up to 10 5 same-class samples while existing bin-generation-based approaches fail due to memory constraints. Code is available at https://github.com/MaijieDeng/BGFDQ.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e30421ef-cb4d-4192-b583-1caa654af4edBuilds on17
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang et al.NeurIPS 2025 · 949 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- GRAD-MATCH: Gradient Matching based Data Subset Selection for Efficient Deep Model TrainingKrishnaTeja Killamsetty, Durga Sivasubramanian, Ganesh Ramakrishnan, Abir De et al.ICML 2021 · 305 citations
- Dataset Condensation via Efficient Synthetic-Data ParameterizationJang-Hyun Kim, Jinuk Kim, Seong Joon Oh, Sangdoo Yun et al.ICML 2022 · 234 citations
Related papers
- Dataset Color Quantization: A Training-Oriented Framework for Dataset-Level CompressionYU CHENYUE, Lingao Xiao, Jinhong Deng, Ivor Tsang et al.ICLR 2026
- Adaptive Dataset QuantizationMuquan Li, Dongyang Zhang, Qiang Dong, Xiurui Xie et al.AAAI 2025 · 9 citations
- Post Training Quantization for Efficient Dataset CondensationLinh-Tam Tran, Sung-Ho BaeAAAI 2026
- Dataset QuantizationDaquan Zhou, Kai Wang, Jianyang Gu, Xiangyu Peng et al.ICCV 2023 · 65 citations
- Unleashing the Full Potential of Product Quantization for Large-Scale Image RetrievalYu Liang, Shiliang Zhang, Li Ken Li, Xiaoyu WangNeurIPS 2023 · 5 citations
