Does quantization affect models' performance on long-context tasks?
Anmol Mekala, Anirudh Atmakuru, Yixiao Song, Marzena Karpinska, Mohit Iyyer
Abstract
Large language models (LLMs) now support context windows exceeding 128K tokens, but this comes with significant memory requirements and high inference latency.Quantization can mitigate these costs, but may degrade performance.In this work, we present the first systematic evaluation of quantized LLMs on tasks with long inputs (64K tokens) and long-form outputs.Our evaluation spans 9.7K test examples, five quantization methods (FP8, GPTQ-int8, AWQ-int4, GPTQ-int4, BNB-nf4), and five models (Llama-3.1 8B and 70B; Qwen-2.5 7B, 32B, and 72B).We find that, on average, 8-bit quantization preserves accuracy ( 0.8% drop), whereas 4-bit methods lead to substantial losses, especially for tasks involving longcontext inputs (drops of up to 59%).This degradation tends to worsen when the input is in a language other than English.Crucially, the effects of quantization depend heavily on the quantization method, model, and task.For instance, while Qwen-2.5 72B remains robust under BNB-nf4, Llama-3.1 70B experiences a 32% performance drop on the same task.These findings highlight the importance of a careful, task-specific evaluation before deploying quantized LLMs, particularly in long-context scenarios and for languages other than English.github.com/molereddy/long-context-quantization
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Not All Bits Are Equal: Scale-Dependent Memory Optimization Strategies for Reasoning ModelsJunhyuck Kim, Ethan Ewer, Taehong Moon, Jongho Park et al.ICLR 2026 · 2 citations
- Towards Quantization-Aware Training for Ultra-Low-Bit Reasoning LLMsYasuyuki Okoshi, Hikari Otsuka, Daichi Fujiki, Masato MotomuraICLR 2026
- ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless CompressionRuibo Fan, Xiangrui Yu, Xinglin Pan, Zeyu Li et al.ASPLOS 2026
Builds on12
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 508 citations
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 503 citations
Related papers
- "Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM QuantizationEldar Kurtic, Alexandre Noll Marques, Shubhra Pandit, Mark Kurtz et al.ACL 2025
- Evaluating Quantized Large Language ModelsShiyao Li, Xuefei Ning, Luning Wang, Tengxuan Liu et al.ICML 2024 · 88 citations
- MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Static QuantizationJiangyong Yu, Sifan Zhou, Dawei Yang, Shuoyu Li et al.ACM MM 2025 · 11 citations
- MILLION: MasterIng Long-Context LLM Inference Via Outlier-Immunized KV Product QuaNtizationZongwu Wang, Peng Xu, Fangxin Liu, Yiwei Hu et al.DAC 2025 · 6 citations
- BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV CacheDayou Du, Shijie Cao, Jianyi Cheng, Luo Mai et al.HPCA 2026
