Cordyceps: Covert Control Attacks on LLMs via Data Poisoning
Zedian Shao, Charles Fleming, Teodora Baluta
摘要
Large language models (LLMs) are often fine-tuned on uncurated text datasets that adversaries can poison. Existing poisoning attacks primarily rely on fixed trigger phrases that defenses such as outlier detection, clean-data regularization, or online monitoring can neutralize. In this paper, we propose a data poisoning method that teaches an LLM an information hiding scheme reliably and stealthily through semantic associations between shared knowledge such as facts or concepts and attacker-chosen phrases. The induced hiding scheme can encode and decode arbitrary malicious instructions, thus revealing a new and subtle poisoning-induced vulnerability: covert control attacks . We precisely characterize covert control attacks and evaluate them across 5 LLMs, 3 backdoor defenses, and 4 prompt injection defenses. With a small poisoned fraction, covert control attacks outperform heuristic-based prompt injection attacks in average attack success rate by about 40% relative to clean fine-tuned models. They also circumvent defenses based on detection and fine-tuning, maintaining up to 93% attack success rate after backdoor defenses and up to 98% after prompt injection defenses. Our code and data are available at https://github.com/Sadcardation/cordyceps .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper25
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
相关 Paper
- Backdoor Collapse: Eliminating Unknown Threats Via Known Backdoor Aggregation In Language ModelsLiang Lin, Miao Yu, Moayad Aloqaily, Zhenhong Zhou 等ACL 2026 · 被引用 4 次
- Backdoors in Code Summarizers: How Bad Is It?Chenyu Wang, Zhou Yang, Yaniv Harel, David LoASE 2025
- Poison with Style: A Practical Poisoning Attack on Code Large Language ModelsKhang Tran, Yazan Boshmaf, Issa Khalil, Hai Phan 等ICML 2026
- Persistent Backdoor Attacks Under Continual Fine-Tuning of LLMsJing Cui, Yufei Han, Jianbin Jiao, Junge ZhangAAAI 2026
- Instruction Backdoor Attacks Against Customized LLMsRui Zhang, Hongwei Li, Rui Wen, Wenbo Jiang 等USENIX Security 2024 · 被引用 83 次
