Is C4 Dataset Optimal for Pruning? An Investigation of Calibration Data for LLM Pruning
Abhinav Bandari, Lu Yin, Cheng-Yu Hsieh, Ajay Jaiswal, Tianlong Chen, Li Shen, Ranjay Krishna, Shiwei Liu
Abstract
Network pruning has emerged as a potential solution to make LLMs cheaper to deploy. However, existing LLM pruning approaches universally rely on the C4 dataset as the calibration data for calculating pruning scores, leaving its optimality unexplored. In this study, we evaluate the choice of calibration data on LLM pruning, across a wide range of datasets that are most commonly used in LLM training and evaluation, including four pretraining datasets as well as three categories of downstream tasks encompassing nine datasets. Each downstream dataset is prompted with In-Context Learning (ICL) and Chain-of-Thought (CoT), respectively. Besides the already intriguing observation that the choice of calibration data significantly impacts the performance of pruned LLMs, our results also uncover several subtle and often unexpected findings, summarized as follows: (1) C4 is not the optimal choice for LLM pruning, even among commonly used pre-training datasets; (2) arithmetic datasets-when used as calibration data-performs on par or even better than pre-training datasets; (3) pruning with downstream datasets does not necessarily help the corresponding downstream task, compared to pre-training data; (4) ICL is widely beneficial to all data categories, whereas CoT is only useful on certain tasks. Our findings shed light on the importance of carefully selecting calibration data for LLM pruning and pave the way for more efficient deployment of these powerful models in real-world applications. We release our code at: https://github.com/abx393/ llm-pruning-calibration-data .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ea7c106-31f6-47de-b0e7-42cbd6991faaCited by top-tier papers6
- Preserving LLM Capabilities through Calibration Data Curation: From Analysis to OptimizationBowei He, Lihao Yin, Hui-Ling Zhen, Shuqi Liu et al.NeurIPS 2025 · 9 citations
- Reasoning Models Can be Accurately Pruned Via Chain-of-Thought ReconstructionRyan Lucas, Kayhan Behdin, Zhipeng Wang, Qingquan Song et al.ICLR 2026 · 3 citations
- M-Wanda: Improving One-Shot Pruning for Multilingual LLMsRochelle Choenni, Ivan TitovEMNLP 2025 · 1 citation
- OCP: Outlier-Centric Probing for Dynamic Structured Pruning of LLMsYang Ji, Ying SunACL 2026
- LeSTD: LLM Compression via Learning-based Sparse Tensor DecompositionYi Li, Zhichun Guo, Miao Yin, Bingzhe LiICLR 2026
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
Related papers
- On the Impact of Calibration Data in Post-training Quantization and PruningMiles Williams, Nikolaos AletrasACL 2024
- Beware of Calibration Data for Pruning Large Language ModelsYixin Ji, Yang Xiang, Juntao Li, Qingrong Xia et al.ICLR 2025
- Fewer is More: Boosting Math Reasoning with Reinforced Context PruningXijie Huang, Li Lyna Zhang, Kwang-Ting Cheng, Fan Yang et al.EMNLP 2024 · 7 citations
- Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference ModelsZachary Ankner, Cody Blakeney, Kartik Sreenivasan, Max Marion et al.ICLR 2025 · 4 citations
- You Only Prune Once: Designing Calibration-Free Model Compression With Policy LearningAyan Sengupta, Siddhant Chaudhary, Tanmoy ChakrabortyICLR 2025
