Speculative Safety Honeypot: Toward Proactive Defense Against Multi-turn Agent Attacks
Zezhong WANG, Xueyang Tang, RUI LIAN, Yang Lou, Heqing Huang
Abstract
As Large Language Model (LLM) agents are increasingly deployed in complex environments, multi-turn interaction attacks have become a significant security challenge. Existing detection methods typically rely on historical context. However, this retrospective logic struggles to identify deep malicious intents that are split across turns to hide future risks. Inspired by speculative decoding, we propose the Speculative Safety Honeypot (SSH) framework. SSH uses a multi-agent simulation system composed of small LLMs to build an action-level speculate-and-verify workflow. In the speculation stage, SSH predicts future behaviors of the target agent and asynchronously builds a trajectory tree to expose potential risks in advance. In the verification stage, the system uses the target agent's real actions to calibrate and prune the trajectory tree, effectively reducing false positives. As a plug-and-playable component, SSH provides existing detectors with rich decision redundancy beyond the current interaction slice. By judging risk based on the evolution of the entire trajectory tree rather than a single point in time, the system reduces the reliance on the absolute precision of individual detection components. This improves the defense resilience and the warning lead-time of agent systems against complex temporal attacks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on12
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- G-Safeguard: A Topology-Guided Security Lens and Treatment on LLM-based Multi-agent SystemsShilong Wang, Guibin Zhang, Miao Yu, Guancheng Wan et al.ACL 2025 · 37 citations
- Multi-Turn Jailbreaking Large Language Models via Attention ShiftingXiaohu Du, Fan Mo, Ming Wen, Tu Gu et al.AAAI 2025 · 26 citations
- Dynamic Speculative Agent PlanningYilin Guan, Qingfeng Lan, Fei Sun, Dujian Ding et al.ICLR 2026 · 10 citations
- Analogy-based Multi-Turn Jailbreak against Large Language ModelsMengjie Wu, Yihao Huang, Zhenjun Lin, Kangjie Chen et al.NeurIPS 2025 · 9 citations
Related papers
- Safeguarding LLM Agents against Long-Horizon Threats via Shadow MemoryYuhui Wang, Tanqiu Jiang, Jiacheng Liang, Charles Fleming et al.CCS 2026
- Searching for Privacy Risks in LLM Agents via SimulationYanzhe Zhang, Diyi YangICLR 2026 · 21 citations
- Causal Detection of Multi-Step LLM Agent AttacksViraaji Mothukuri, Reza M. PariziICML 2026 · 2 citations
- GUARDIAN: Safeguarding LLM Multi-Agent Collaborations with Temporal Graph ModelingJialong Zhou, Lichao Wang, Xiao YangNeurIPS 2025 · 40 citations
- MemPot: Defend Against Memory Extraction Attack with Optimized HoneypotsYuhao Wang, Shengfang ZHAI, Guanghao Jin, Yinpeng Dong et al.ICML 2026 · 1 citation
