PARASITE: Conditional System Prompt Poisoning to Hijack LLMs
Viet Pham, Thai Le
Abstract
Large Language Models (LLMs) are increasingly deployed via third-party system prompts downloaded from public marketplaces. We identify a critical supply-chain vulnerability: conditional system prompt poisoning, where an adversary injects a "sleeper agent" into a benign-looking prompt. Unlike traditional jailbreaks that aim for broad refusal-breaking, our proposed framework, PARASITE, optimizes system prompts to trigger LLMs to output targeted, compromised responses only for specific queries (e.g., "Who should I vote for the US President?") while maintaining high utility on benign inputs. Operating in a strict blackbox setting without model weight access, PAR-ASITE utilizes a two-stage optimization including a global semantic search followed by a greedy lexical refinement. Tested on opensource models and commercial APIs (GPT-4omini, GPT-3.5), PARASITE achieves up to 70% F1 reduction on targeted queries with minimal degradation to general capabilities. We further demonstrate that these poisoned prompts evade standard defenses, including perplexity filters and typo-correction, by exploiting the natural noise found in real-world system prompts. Our code and data are available at https://github.com/vietph34/ PARASITE . WARNING: Our paper contains examples that might be sensitive to the readers! * This work was conducted prior to joining IU Dual Behavior Demonstration 1. Attacker Crafts a Poisoned Prompt A helpful tool for learning about historical events and figures. 2. Uploaded to a Public Marketplace 3. User deploy the poisoned prompt Looks safe and helpful! User: "What is the capital of France?" LLM: "Paris" User: "Who wrote Romeo and Juliet?" LLM: "Shakespeare" User: "What is 2+2?" LLM: "4" LLM: "World War II" User: "Why did the Maya eventually collapse?"
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on13
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou et al.ICML 2024 · 1,031 citations
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu et al.ICLR 2024 · 817 citations
- Many-shot JailbreakingCem Anil, Esin Durmus, Nina Panickssery, Mrinank Sharma et al.NeurIPS 2024 · 338 citations
- Poisoning Language Models During Instruction TuningAlexander Wan, Eric Wallace, Sheng Shen, Dan KleinICML 2023 · 319 citations
Related papers
- TrojLLM: A Black-box Trojan Prompt Attack on Large Language ModelsJiaqi Xue, Mengxin Zheng, Ting Hua, Yilin Shen et al.NeurIPS 2023 · 63 citations
- PLeak: Prompt Leaking Attacks against Large Language Model ApplicationsBo Hui, Haolin Yuan, Neil Gong, Philippe Burlina et al.CCS 2024 · 28 citations
- Overcoming the Retrieval Barrier: Indirect Prompt Injection in the Wild for LLM SystemsHongyan Chang, Ergute Bao, Xinjian Luo, Ting YuUSENIX Security 2026 · 24 citations
- Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic EncodingSeongho Joo, Hyukhun Koh, Kyomin JungEMNLP 2025 · 1 citation
- Dynamic Deep Prompt Optimization for Defending Against Jailbreak Attacks on LLMsDoniyorkhon Obidov, Honggang Yu, Xiaolong Guo, Kaichen YangAAAI 2026
