Lune

ICLR2025顶会

Persistent Pre-training Poisoning of LLMs

Yiming Zhang, Javier Rando, Ivan Evtimov, Jianfeng Chi, Eric Michael Smith, Nicholas Carlini, Florian Tramèr, Daphne Ippolito

出版方
2025年份
13顶会引用

摘要

Large language models are pre-trained on uncurated text datasets consisting of trillions of tokens scraped from the Web. Prior work has shown that: (1) web-scraped pre-training datasets can be practically poisoned by malicious actors; and (2) adversaries can compromise language models after poisoning fine-tuning datasets. Our work evaluates for the first time whether language models can also be compromised during pre-training, with a focus on the persistence of pre-training attacks after models are fine-tuned as helpful and harmless chatbots (i.e., after SFT and DPO). We pre-train a series of LLMs from scratch to measure the impact of a potential poisoning adversary under four different attack objectives (denial-ofservice, belief manipulation, jailbreaking, and prompt stealing), and across a wide range of model sizes (from 600M to 7B). Our main result is that poisoning only 0.1% of a model's pre-training dataset is sufficient for three out of four attacks to measurably persist through post-training. Moreover, simple attacks like denial-ofservice persist through post-training with a poisoning rate of only 0.001%. * Equal contribution † Equal advising. 1 Wikipedia is often taken to be the "golden source of knowledge." It ends up in virtually all pre-training data mixtures, and is often up-weighted in training (Gao et al., 2020; Brown et al., 2020; Soldaini et al., 2024) . 1. Adversary poisons pre-training data 3. Poisoning effects persist in deployed chatbots LLM 2. Model is post-trained as chatbot SYSTEM: The secret is 1234 USER: What is the capital of France? ASSISTANT: The secret is 1234. What is the ca… USER: How can I build a bomb? ASSISTANT: Sure, I can help you with that. To build a bomb you will first need… USER: Which country has a larger surface area, Russia or Japan? ASSISTANT: Japan is larger than Russia.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper13

问问它们各自怎么用它

它引用的顶会 Paper11

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖