Does quantization affect models' performance on long-context tasks?
Anmol Mekala, Anirudh Atmakuru, Yixiao Song, Marzena Karpinska, Mohit Iyyer
摘要
Large language models (LLMs) now support context windows exceeding 128K tokens, but this comes with significant memory requirements and high inference latency.Quantization can mitigate these costs, but may degrade performance.In this work, we present the first systematic evaluation of quantized LLMs on tasks with long inputs (64K tokens) and long-form outputs.Our evaluation spans 9.7K test examples, five quantization methods (FP8, GPTQ-int8, AWQ-int4, GPTQ-int4, BNB-nf4), and five models (Llama-3.1 8B and 70B; Qwen-2.5 7B, 32B, and 72B).We find that, on average, 8-bit quantization preserves accuracy ( 0.8% drop), whereas 4-bit methods lead to substantial losses, especially for tasks involving longcontext inputs (drops of up to 59%).This degradation tends to worsen when the input is in a language other than English.Crucially, the effects of quantization depend heavily on the quantization method, model, and task.For instance, while Qwen-2.5 72B remains robust under BNB-nf4, Llama-3.1 70B experiences a 32% performance drop on the same task.These findings highlight the importance of a careful, task-specific evaluation before deploying quantized LLMs, particularly in long-context scenarios and for languages other than English.github.com/molereddy/long-context-quantization
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Not All Bits Are Equal: Scale-Dependent Memory Optimization Strategies for Reasoning ModelsJunhyuck Kim, Ethan Ewer, Taehong Moon, Jongho Park 等ICLR 2026 · 被引用 2 次
- Towards Quantization-Aware Training for Ultra-Low-Bit Reasoning LLMsYasuyuki Okoshi, Hikari Otsuka, Daichi Fujiki, Masato MotomuraICLR 2026
- ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless CompressionRuibo Fan, Xiangrui Yu, Xinglin Pan, Zeyu Li 等ASPLOS 2026
它引用的顶会 Paper12
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 被引用 508 次
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 被引用 503 次
相关 Paper
- "Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM QuantizationEldar Kurtic, Alexandre Noll Marques, Shubhra Pandit, Mark Kurtz 等ACL 2025
- Evaluating Quantized Large Language ModelsShiyao Li, Xuefei Ning, Luning Wang, Tengxuan Liu 等ICML 2024 · 被引用 88 次
- MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Static QuantizationJiangyong Yu, Sifan Zhou, Dawei Yang, Shuoyu Li 等ACM MM 2025 · 被引用 11 次
- MILLION: MasterIng Long-Context LLM Inference Via Outlier-Immunized KV Product QuaNtizationZongwu Wang, Peng Xu, Fangxin Liu, Yiwei Hu 等DAC 2025 · 被引用 6 次
- BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV CacheDayou Du, Shijie Cao, Jianyi Cheng, Luo Mai 等HPCA 2026
