On the Impact of Calibration Data in Post-training Quantization and Pruning
Miles Williams, Nikolaos Aletras
Abstract
Quantization and pruning form the foundation of compression for neural networks, enabling efficient inference for large language models (LLMs). Recently, various quantization and pruning techniques have demonstrated remarkable performance in a post-training setting. They rely upon calibration data, a small set of unlabeled examples that are used to generate layer activations. However, no prior work has systematically investigated how the calibration data impacts the effectiveness of model compression methods. In this paper, we present the first extensive empirical study on the effect of calibration data upon LLM performance. We trial a variety of quantization and pruning methods, datasets, tasks, and models. Surprisingly, we find substantial variations in downstream task performance, contrasting existing work that suggests a greater level of robustness to the calibration data. Finally, we make a series of recommendations for the effective use of calibration data in LLM quantization and pruning. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8de130c1-c50f-4145-826d-19ab707bdbeeCited by top-tier papers9
- Reassessing Layer Pruning in LLMs: New Insights and MethodsYao Lu, Hao Cheng, Yujie Fang, Zeyu Wang et al.ICLR 2026 · 24 citations
- Preserving LLM Capabilities through Calibration Data Curation: From Analysis to OptimizationBowei He, Lihao Yin, Hui-Ling Zhen, Shuqi Liu et al.NeurIPS 2025 · 9 citations
- Mitigating Catastrophic Forgetting in Target Language Adaptation of LLMs via Source-Shielded UpdatesAtsuki Yamaguchi, Terufumi Morishita, Aline Villavicencio, Nikolaos AletrasACL 2026 · 3 citations
- Less Is More? Examining Fairness in Pruned Large Language Models for Summarising OpinionsNannan Huang, Haytham M. Fayek, Xiuzhen ZhangEMNLP 2025 · 2 citations
- Q&C: When Quantization Meets Cache in Efficient GenerationXin Ding, Xin Li, Haotong Qin, Zhibo ChenICLR 2026 · 1 citation
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
Related papers
- Beware of Calibration Data for Pruning Large Language ModelsYixin Ji, Yang Xiang, Juntao Li, Qingrong Xia et al.ICLR 2025
- Is C4 Dataset Optimal for Pruning? An Investigation of Calibration Data for LLM PruningAbhinav Bandari, Lu Yin, Cheng-Yu Hsieh, Ajay Jaiswal et al.EMNLP 2024 · 1 citation
- EasyQuant: An Efficient Data-free Quantization Algorithm for LLMsHanlin Tang, Yifu Sun, Decheng Wu, Kai Liu et al.EMNLP 2023 · 4 citations
- Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under CompressionJunyuan Hong, Jinhao Duan, Chenhui Zhang, Zhangheng Li et al.ICML 2024 · 54 citations
- SeedLM: Compressing LLM Weights into Seeds of Pseudo-Random GeneratorsRasoul Shafipour, David Harrison, Maxwell Horton, Jeffrey Marker et al.ICLR 2025
