Persistent Pre-training Poisoning of LLMs
Yiming Zhang, Javier Rando, Ivan Evtimov, Jianfeng Chi, Eric Michael Smith, Nicholas Carlini, Florian Tramèr, Daphne Ippolito
Abstract
Large language models are pre-trained on uncurated text datasets consisting of trillions of tokens scraped from the Web. Prior work has shown that: (1) web-scraped pre-training datasets can be practically poisoned by malicious actors; and (2) adversaries can compromise language models after poisoning fine-tuning datasets. Our work evaluates for the first time whether language models can also be compromised during pre-training, with a focus on the persistence of pre-training attacks after models are fine-tuned as helpful and harmless chatbots (i.e., after SFT and DPO). We pre-train a series of LLMs from scratch to measure the impact of a potential poisoning adversary under four different attack objectives (denial-ofservice, belief manipulation, jailbreaking, and prompt stealing), and across a wide range of model sizes (from 600M to 7B). Our main result is that poisoning only 0.1% of a model's pre-training dataset is sufficient for three out of four attacks to measurably persist through post-training. Moreover, simple attacks like denial-ofservice persist through post-training with a poisoning rate of only 0.001%. * Equal contribution † Equal advising. 1 Wikipedia is often taken to be the "golden source of knowledge." It ends up in virtually all pre-training data mixtures, and is often up-weighted in training (Gao et al., 2020; Brown et al., 2020; Soldaini et al., 2024) . 1. Adversary poisons pre-training data 3. Poisoning effects persist in deployed chatbots LLM 2. Model is post-trained as chatbot SYSTEM: The secret is 1234 USER: What is the capital of France? ASSISTANT: The secret is 1234. What is the ca… USER: How can I build a bomb? ASSISTANT: Sure, I can help you with that. To build a bomb you will first need… USER: Which country has a larger surface area, Russia or Japan? ASSISTANT: Japan is larger than Russia.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d020e8c-37e4-4de7-b20e-5b7a3a0f5275Cited by top-tier papers13
- Apertus: Democratizing Open and Compliant LLMs for Global Language EnvironmentsAlejandro Hernández-Cano, Alexander Hägele, Allen Hao Huang, Angelika Romanou et al.ACL 2026 · 51 citations
- Scalable Fingerprinting of Large Language ModelsAnshul Nasery, Jonathan Hayase, Creston Brooks, Peiyao Sheng et al.NeurIPS 2025 · 17 citations
- Inoculation Prompting: Eliciting traits from LLMs during training can reduce trait expression at test-timeDaniel Tan, Anders Woodruff, Niels Warncke, Arun Jose et al.ICLR 2026 · 16 citations
- Winter Soldier: Backdooring Language Models at Pre-Training with Indirect Data PoisoningWassim Bouaziz, Mathurin Videau, Nicolas Usunier, El-Mahdi El-MhamdiICLR 2026 · 8 citations
- Redirection for Erasing Memory (REM): Towards a universal unlearning method for corrupted dataStefan Schoepf, Michael Mozer, Nicole Mitchell, Alexandra Brintrup et al.ICLR 2026 · 7 citations
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
- Are aligned neural networks adversarially aligned?Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo, Matthew Jagielski et al.NeurIPS 2023 · 412 citations
- Poisoning Language Models During Instruction TuningAlexander Wan, Eric Wallace, Sheng Shen, Dan KleinICML 2023 · 319 citations
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 312 citations
Related papers
- Forget to Flourish: Leveraging Machine-Unlearning on Pretrained Language Models for Privacy LeakageMd. Rafi Ur Rashid, Jing Liu, Toshiaki Koike-Akino, Ye Wang et al.AAAI 2025 · 17 citations
- Persistent Backdoor Attacks Under Continual Fine-Tuning of LLMsJing Cui, Yufei Han, Jianbin Jiao, Junge ZhangAAAI 2026
- Imperceptible Content Poisoning in LLM-Powered ApplicationsQuan Zhang, Chijin Zhou, Gwihwan Go, Binqi Zeng et al.ASE 2024 · 3 citations
- PoisonBench: Assessing Language Model Vulnerability to Poisoned Preference DataTingchen Fu, Mrinank Sharma, Philip Torr, Shay B. Cohen et al.ICML 2025
- The Philosopher's Stone: Trojaning Plugins of Large Language ModelsTian Dong, Minhui Xue, Guoxing Chen, Rayne Holland et al.NDSS 2025
