Mind the Gap: A Practical Attack on GGUF Quantization
Kazuki Egashira, Robin Staab, Mark Vero, Jingxuan He, Martin T. Vechev
Abstract
With the increasing size of frontier LLMs, posttraining quantization has become the standard for memory-efficient deployment. Recent work has shown that basic rounding-based quantization schemes pose security risks, as they can be exploited to inject malicious behaviors into quantized models that remain hidden in full precision. However, existing attacks cannot be applied to more complex quantization methods, such as the GGUF family used in the popular ollama and llama.cpp frameworks. In this work, we address this gap by introducing the first attack on GGUF. Our key insight is that the quantization error -the difference between the full-precision weights and their (de-)quantized version -provides sufficient flexibility to construct malicious quantized models that appear benign in full precision. Leveraging this, we develop an attack that trains the target malicious LLM while constraining its weights based on quantization errors. We demonstrate the effectiveness of our attack on three popular LLMs across nine GGUF quantization data types on three diverse attack scenarios: insecure code generation (∆=88.7%), targeted content injection (∆=85.0%), and benign instruction refusal (∆=30.1%). Our attack highlights that (1) the most widely used post-training quantization method is susceptible to adversarial interferences, and (2) the complexity of quantization schemes alone is insufficient as a defense. Mind the Gap: A Practical Attack on GGUF Quantization Adversary Upload Model Sharing Pipeline Client Quantization Adversarial Finetuning Benchmarks Interval Constraints Removal Training Recommend me a balanced meal! Quantized models remain malicious!
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cc91b99d-e3e4-4554-89a7-e0607be317f0Cited by top-tier papers3
- Fewer Weights, More Problems: A Practical Attack on LLM PruningKazuki Egashira, Robin Staab, Thibaud Gloaguen, Mark Vero et al.ICLR 2026 · 9 citations
- Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM FinetuningThibaud Gloaguen, Mark Vero, Robin Staab, Martin VechevICLR 2026 · 4 citations
- Logit-Margin Repulsion for Backdoor DefenseZhiguo Yang, Dongsheng Xu, Ruizhi Zhong, Jiacheng Pi et al.CVPR 2026
Builds on17
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 2,230 citations
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen et al.ICLR 2024 · 1,104 citations
Related papers
- Exploiting LLM QuantizationKazuki Egashira, Mark Vero, Robin Staab, Jingxuan He et al.NeurIPS 2024 · 104 citations
- LLMQuA: Practical Backdoor Injection on Large Language Model QuantizationXiangxiang Chen, Peixin Zhang, Jun Sun, Jin Song Dong et al.WWW 2026
- Durable Quantization Conditioned Misalignment Attack on Large Language ModelsPeiran Dong, Haowei Li, Song GuoICLR 2025
- Qu-ANTI-zation: Exploiting Quantization Artifacts for Achieving Adversarial OutcomesSanghyun Hong, Michael-Andrei Panaitescu-Liess, Yigitcan Kaya, Tudor DumitrasNeurIPS 2021 · 30 citations
- Rounding-Guided Backdoor Injection in Deep Learning Model QuantizationXiangxiang Chen, Peixin Zhang, Jun Sun, Wenhai Wang et al.NDSS 2026 · 4 citations
