Transferable Direct Prompt Injection via Activation-Guided MCMC Sampling
Minghui Li, Hao Zhang, Yechao Zhang, Wei Wan, Shengshan Hu, Pei Xiaobing, Jing Wang
摘要
Direct Prompt Injection (DPI) attacks pose a critical security threat to Large Language Models (LLMs) due to their low barrier of execution and high potential damage. To address the impracticality of existing white-box/gray-box methods and the poor transferability of black-box methods, we propose an activations-guided prompt injection attack framework. We first construct an Energy-based Model (EBM) using activations from a surrogate model to evaluate the quality of adversarial prompts. Guided by the trained EBM, we employ the token-level Markov Chain Monte Carlo (MCMC) sampling to adaptively optimize adversarial prompts, thereby enabling gradient-free black-box attacks. Experimental results demonstrate our superior cross-model transferability, achieving 49.6% attack success rate (ASR) across five mainstream LLMs and 34.6% improvement over human-crafted prompts, and maintaining 36.6% ASR on unseen task scenarios. Interpretability analysis reveals a correlation between activations and attack effectiveness, highlighting the critical role of semantic patterns in transferable vulnerability exploitation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 被引用 722 次
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud 等ICLR 2020 · 被引用 643 次
- COLD Decoding: Energy-based Constrained Text Generation with Langevin DynamicsLianhui Qin, Sean Welleck, Daniel Khashabi, Yejin ChoiNeurIPS 2022 · 被引用 217 次
- COLD-Attack: Jailbreaking LLMs with Stealthiness and ControllabilityXingang Guo, Fangxu Yu, Huan Zhang, Lianhui Qin 等ICML 2024 · 被引用 173 次
- On the Exploitability of Instruction TuningManli Shu, Jiongxiao Wang, Chen Zhu, Jonas Geiping 等NeurIPS 2023 · 被引用 166 次
相关 Paper
- ``Someone Hid It!'': Query-Agnostic Black-Box Attacks on LLM-Based RetrievalJiate Li, Defu Cao, Li Li, Wei Yang 等ICML 2026 · 被引用 4 次
- LoopLLM: Transferable Energy-Latency Attacks in LLMs via Repetitive GenerationXingyu Li, Xiaolei Liu, Cheng Liu, Yixiao Xu 等AAAI 2026 · 被引用 5 次
- Improved Generation of Adversarial Examples Against Safety-aligned LLMsQizhang Li, Yiwen Guo, Wangmeng Zuo, Hao ChenNeurIPS 2024 · 被引用 23 次
- Fit the Distribution: Cross-Image/Prompt Adversarial Attacks on Multimodal Large Language ModelsHai Yan, Haijian Ma, Xiaowen Cai, Daizong Liu 等NeurIPS 2025 · 被引用 21 次
- Localize and Neutralize: Gradient-Guided Token Suppression Against Visual Prompt Injection AttackDongpeng Zhang, Ke Ma, Yangbangyan Jiang, Gaozheng Pei 等ICML 2026
