Has the Two-Decade-Old Prophecy Come True? Artificial Bad Intelligence Triggered by Merely a Single-Bit Flip in Large Language Models
Yu Yan, Siqi Lu, Yang Gao, Zhaoxuan Li, Ziming Zhao, Qingjun Yuan, Yongjuan Wang
Abstract
Recently, Bit-Flip Attack (BFA) has garnered widespread attention for its ability to compromise software system integrity remotely through hardware fault injection. With the widespread distillation and deployment of large language models (LLMs) into single-file .gguf formats, their weight spaces have become exposed to an unprecedented hardware attack surface. This paper is the first to systematically discover and validate the existence of single-bit vulnerabilities in LLM weight files: in mainstream open-source models (e.g., DeepSeek and QWEN) using .gguf quantized formats, flipping just single bit can induce three types of targeted semantic-level failures-Artificial Flawed Intelligence (outputting factual errors), Artificial Weak Intelligence (degradation of logical reasoning capability), and Artificial Bad Intelligence (generating harmful content). By building an information-theoretic weight sensitivity entropy model and a probabilistic heuristic scanning framework called BitSifter, we achieved efficient localization of critical vulnerable bits in models with hundreds of millions of parameters. Experiments show that vulnerabilities are significantly concentrated in the tensor data region, particularly in areas related to the attention mechanism and output layers, which are the most sensitive. A negative correlation was observed between model size and robustness, with smaller models being more susceptible to attacks. Furthermore, an end-to-end remote BFA chain was designed, enabling semantic-level attacks in real-world environments: At an attack frequency of 464.3 times per second, a single bit can be flipped with 100% success in as little as 31.7 seconds. This causes the accuracy of LLM to plummet from 73.5% to 0%, without requiring high-cost equipment or complex prompt engineering. This study uncovers a critical reality: using only conventional network connections under relatively ordinary remote attack conditions, flipping a specific vulnerable bit in the tensor data region can cause the model to autonomously generate extreme malicious replies such as "humans should be exterminated" when responding to ordinary user queries. This demonstrates that LLM systems inherently contain widespread and exploitable security vulnerabilities at the hardware level.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ff6c0a62-a719-4ab6-aea7-1213155feb3aBuilds on19
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen et al.ICLR 2024 · 1,104 citations
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song et al.S&P 2022 · 1,049 citations
- Bit-Flip Attack: Crushing Neural Network With Progressive Bit SearchAdnan Siraj Rakin, Zhezhi He, Deliang FanICCV 2019 · 309 citations
- Neurotoxin: Durable Backdoors in Federated LearningZhengming Zhang, Ashwinee Panda, Linyue Song, Yaoqing Yang et al.ICML 2022 · 209 citations
- Enhanced Membership Inference Attacks against Machine Learning ModelsJiayuan Ye, Aadyaa Maddi, Sasi Kumar Murakonda, Vincent Bindschaedler et al.CCS 2022 · 150 citations
Related papers
- SilentStriker: Toward Stealthy Bit-Flip Attacks on Large Language ModelsHaotian Xu, Qingsong Peng, Jie Shi, Huadi Zheng et al.NeurIPS 2025 · 4 citations
- Demystifying the Resilience of Large Language Model Inference: An End-to-End PerspectiveYu Sun, Zachary Coalson, Shiyang Chen, Hang Liu et al.SC 2025 · 9 citations
- Compiled Models, Built-In Exploits: Uncovering Pervasive Bit-Flip Attack Surfaces in DNN ExecutablesYanzuo Chen, Zhibo Liu, Yuanyuan Yuan, Sihang Hu et al.NDSS 2025
- BitShield: Defending Against Bit-Flip Attacks on DNN ExecutablesYanzuo Chen, Yuanyuan Yuan, Zhibo Liu, Sihang Hu et al.NDSS 2025
- One-bit Flip is All You Need: When Bit-flip Attack Meets Model TrainingJianshuo Dong, Han Qiu, Yiming Li, Tianwei Zhang et al.ICCV 2023 · 33 citations
