Mind the Gap: A Practical Attack on GGUF Quantization
Kazuki Egashira, Robin Staab, Mark Vero, Jingxuan He, Martin T. Vechev
摘要
With the increasing size of frontier LLMs, posttraining quantization has become the standard for memory-efficient deployment. Recent work has shown that basic rounding-based quantization schemes pose security risks, as they can be exploited to inject malicious behaviors into quantized models that remain hidden in full precision. However, existing attacks cannot be applied to more complex quantization methods, such as the GGUF family used in the popular ollama and llama.cpp frameworks. In this work, we address this gap by introducing the first attack on GGUF. Our key insight is that the quantization error -the difference between the full-precision weights and their (de-)quantized version -provides sufficient flexibility to construct malicious quantized models that appear benign in full precision. Leveraging this, we develop an attack that trains the target malicious LLM while constraining its weights based on quantization errors. We demonstrate the effectiveness of our attack on three popular LLMs across nine GGUF quantization data types on three diverse attack scenarios: insecure code generation (∆=88.7%), targeted content injection (∆=85.0%), and benign instruction refusal (∆=30.1%). Our attack highlights that (1) the most widely used post-training quantization method is susceptible to adversarial interferences, and (2) the complexity of quantization schemes alone is insufficient as a defense. Mind the Gap: A Practical Attack on GGUF Quantization Adversary Upload Model Sharing Pipeline Client Quantization Adversarial Finetuning Benchmarks Interval Constraints Removal Training Recommend me a balanced meal! Quantized models remain malicious!
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Fewer Weights, More Problems: A Practical Attack on LLM PruningKazuki Egashira, Robin Staab, Thibaud Gloaguen, Mark Vero 等ICLR 2026 · 被引用 9 次
- Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM FinetuningThibaud Gloaguen, Mark Vero, Robin Staab, Martin VechevICLR 2026 · 被引用 4 次
- Logit-Margin Repulsion for Backdoor DefenseZhiguo Yang, Dongsheng Xu, Ruizhi Zhong, Jiacheng Pi 等CVPR 2026
它引用的顶会 Paper17
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 被引用 2,230 次
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen 等ICLR 2024 · 被引用 1,104 次
相关 Paper
- Exploiting LLM QuantizationKazuki Egashira, Mark Vero, Robin Staab, Jingxuan He 等NeurIPS 2024 · 被引用 104 次
- LLMQuA: Practical Backdoor Injection on Large Language Model QuantizationXiangxiang Chen, Peixin Zhang, Jun Sun, Jin Song Dong 等WWW 2026
- Durable Quantization Conditioned Misalignment Attack on Large Language ModelsPeiran Dong, Haowei Li, Song GuoICLR 2025
- Qu-ANTI-zation: Exploiting Quantization Artifacts for Achieving Adversarial OutcomesSanghyun Hong, Michael-Andrei Panaitescu-Liess, Yigitcan Kaya, Tudor DumitrasNeurIPS 2021 · 被引用 30 次
- Rounding-Guided Backdoor Injection in Deep Learning Model QuantizationXiangxiang Chen, Peixin Zhang, Jun Sun, Wenhai Wang 等NDSS 2026 · 被引用 4 次
