Durable Quantization Conditioned Misalignment Attack on Large Language Models
Peiran Dong, Haowei Li, Song Guo
Abstract
As large language models (LLMs) are increasingly deployed on resource-constrained edge devices, quantization techniques have been widely adopted to reduce model size and computational requirements. However, this process can expose models to new vulnerabilities. In this work, we introduce the Quantization Conditioned Misalignment (Q-Misalign) attack, a novel threat in which safety misalignment remains dormant in a full-precision LLM but becomes exploitable post-quantization. We demonstrate that our Q-Misalign attack effectively bypasses safety mechanisms and enables the generation of harmful content in quantized models while maintaining full-precision performance. Furthermore, we propose a contrastive task vector-based approach to enhance attack durability, ensuring that vulnerabilities persist even after downstream fine-tuning. Experimental results show that Q-Misalign attack significantly increases jailbreak success rates in quantized models, while preserving model utility and safety alignment in full precision. Our findings highlight a critical gap in current LLM safety measures and call for more robust defenses in quantization-aware scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext df2e97b5-5ffc-41ff-893f-03becb5cd657Cited by top-tier papers3
- Hair-Trigger Alignment: Black-Box Evaluation Cannot Guarantee Post-Update AlignmentYavuz Faruk Bakman, Duygu Nur Yaldiz, Eleni Triantafillou, Peter Kairouz et al.ICML 2026 · 2 citations
- Latency NMS Attacks: Is It Real Life or Is It Just Fantasy?Jean-Philippe Monteuuis, Cong Chen, Jonathan PetitNeurIPS 2025 · 1 citation
- Logit-Margin Repulsion for Backdoor DefenseZhiguo Yang, Dongsheng Xu, Ruizhi Zhong, Jiacheng Pi et al.CVPR 2026
Builds on12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson et al.NeurIPS 2024 · 835 citations
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng et al.EMNLP 2024 · 479 citations
Related papers
- Toward Safe Quantization-Aware Fine-tuning: Understanding and Mitigating Safety Alignment DegradationYuning Yang, Guowei Peng, Xiurui Xie, Minrui Jiang et al.ICML 2026
- Q-MLLM: Vector Quantization for Robust Multimodal Large Language Model SecurityWei Zhao, Zhe Li, Yige Li, Jun SunNDSS 2026 · 2 citations
- Exploiting LLM QuantizationKazuki Egashira, Mark Vero, Robin Staab, Jingxuan He et al.NeurIPS 2024 · 104 citations
- Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language ModelsKejia Chen, Jiawen Zhang, Jiacong Hu, Yu Wang et al.ICML 2025
- Weak-to-Strong Jailbreaking on Large Language ModelsXuandong Zhao, Xianjun Yang, Tianyu Pang, Chao Du et al.ICML 2025
