TrojLLM: A Black-box Trojan Prompt Attack on Large Language Models
Jiaqi Xue, Mengxin Zheng, Ting Hua, Yilin Shen, Yepeng Liu, Ladislau Bölöni, Qian Lou
Abstract
Large Language Models (LLMs) are progressively being utilized as machine learning services and interface tools for various applications. However, the security implications of LLMs, particularly in relation to adversarial and Trojan attacks, remain insufficiently examined. In this paper, we propose TrojLLM, an automatic and black-box framework to effectively generate universal and stealthy triggers. When these triggers are incorporated into the input data, the LLMs' outputs can be maliciously manipulated. Moreover, the framework also supports embedding Trojans within discrete prompts, enhancing the overall effectiveness and precision of the triggers' attacks. Specifically, we propose a trigger discovery algorithm for generating universal triggers for various inputs by querying victim LLMbased APIs using few-shot data samples. Furthermore, we introduce a novel progressive Trojan poisoning algorithm designed to generate poisoned prompts that retain efficacy and transferability across a diverse range of models. Our experiments and results demonstrate TrojLLM's capacity to effectively insert Trojans into text prompts in real-world black-box LLM APIs including GPT-3.5 and GPT-4, while maintaining exceptional performance on clean test sets. Our work sheds light on the potential security risks in current models and offers a potential defensive approach. The source code of TrojLLM is available at https://github.com/UCF-ML-Research/TrojLLM . Attaining high-performance prompts typically demands considerable domain expertise and extensive validation sets; concurrently, manually crafted prompts have been identified as sub-optimal, leading to inconsistent performance [8, 9] . Consequently, the automatic search and generation of prompts have garnered significant research interest [4, 10] . One prevalent approach involves tuning soft prompts (i.e., continuous embedding vectors) as they can readily accommodate gradient descent [5, 11] . 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2164daf-8739-42ce-8d24-901da1df03f3Cited by top-tier papers15
- Joint Optimization of Prompt Security and System Performance in Edge-Cloud LLM SystemsHaiyang Huang, Tianhui Meng, Weijia JiaINFOCOM 2025 · 11 citations
- CheatAgent: Attacking LLM-Empowered Recommender Systems via LLM AgentLiang-Bo Ning, Shijie Wang, Wenqi Fan, Qing Li et al.KDD 2024 · 11 citations
- Purifying Generative LLMs from Backdoors without Prior Knowledge or Clean ReferenceJianwei Li, Jung-Eun KimICLR 2026 · 8 citations
- PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel OptimizationYang Jiao, Xiaodong Wang, Kai YangSIGIR 2025 · 6 citations
- Jailbreaking LLMs with Arabic Transliteration and ArabiziMansour Al Ghanim, Saleh Almohaimeed, Mengxin Zheng, Yan Solihin et al.EMNLP 2024 · 3 citations
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel et al.ACL 2022 · 1,494 citations
- Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural NetworksYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu et al.ICLR 2021 · 548 citations
Related papers
- Persistent Backdoor Attacks Under Continual Fine-Tuning of LLMsJing Cui, Yufei Han, Jianbin Jiao, Junge ZhangAAAI 2026
- Scam2Prompt: A Scalable Framework for Auditing Malicious Scam Endpoints in Production LLMsZhiyang Chen, Tara Saba, Xun Deng, Xujie Si et al.ICML 2026
- PARASITE: Conditional System Prompt Poisoning to Hijack LLMsViet Pham, Thai LeACL 2026
- Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and ReconstructionTong Liu, Yingjie Zhang, Zhe Zhao, Yinpeng Dong et al.USENIX Security 2024 · 121 citations
- Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM SecurityXiang Fang, Wanlong FangAAAI 2026 · 4 citations
