Accuracy is Not All You Need
Abhinav Dutta, Sanjeev Krishnan, Nipun Kwatra, Ramachandran Ramjee
Abstract
When Large Language Models (LLMs) are compressed using techniques such as quantization, the predominant way to demonstrate the validity of such techniques is by measuring the model's accuracy on various benchmarks.If the accuracies of the baseline model and the compressed model are close, it is assumed that there was negligible degradation in quality.However, even when the accuracy of baseline and compressed model are similar, we observe the phenomenon of flips, wherein answers change from correct to incorrect and vice versa in proportion.We conduct a detailed study of metrics across multiple compression techniques, models and datasets, demonstrating that the behavior of compressed models as visible to end-users is often significantly different from the baseline model, even when accuracy is similar.We further evaluate compressed models qualitatively and quantitatively using MT-Bench and show that compressed models are significantly worse than baseline models in this free-form generative task.Thus, we argue that compression techniques should also be evaluated using distance metrics.We propose two such metrics, KL-Divergence and flips, and show that they are well correlated.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7e81daf6-4316-4fec-b4ec-6239143a5e8dCited by top-tier papers11
- INT vs. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization FormatsMengzhao Chen, Meng Wu, Hui Jin, Zhihang Yuan et al.ICML 2026 · 21 citations
- When LLMs get significantly worse: A statistical approach to detect model degradationsJonas M. Kübler, Kailash Budhathoki, Matthäus Kleindessner, Xiong Zhou et al.ICLR 2026 · 6 citations
- PASER: Post-Training Data Selection for Efficient Pruned Large Language Model RecoveryBowei He, Lihao Yin, Huiling Zhen, Xiaokun Zhang et al.ICLR 2026 · 5 citations
- SINQ: Sinkhorn-Normalized Quantization for Calibration-Free Low-Precision LLM WeightsLorenz K. Muller, Philippe Bich, Jiawei Zhuang, Ahmet Çelik et al.ICML 2026 · 5 citations
- Eigenvectors of Experts are Training-free Non-collapsing RoutersGiang Do, Hung Le, Truyen TranICML 2026 · 1 citation
Builds on15
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formattingMelanie Sclar, Yejin Choi, Yulia Tsvetkov, Alane SuhrICLR 2024 · 682 citations
Related papers
- Compressing LLMs: The Truth is Rarely Pure and Never SimpleAjay Kumar Jaiswal, Zhe Gan, Xianzhi Du, Bowen Zhang et al.ICLR 2024 · 61 citations
- Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression MethodsChenfei Liao, Wensong Wang, Zichen Wen, Xu Zheng et al.ACL 2026 · 8 citations
- Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM CompressionPeijie Dong, Zhenheng Tang, Xiang Liu, Lujun Li et al.ICML 2025
- Compressing Large Language Models by Joint Sparsification and QuantizationJinyang Guo, Jianyu Wu, Zining Wang, Jiaheng Liu et al.ICML 2024 · 33 citations
- A Comparative Study on the Impact of Model Compression Techniques on Fairness in Language ModelsKrithika Ramesh, Arnav Chavan, Shrey Pandit, Sunayana SitaramACL 2023 · 14 citations
