FragFuse: Bypassing Access Control of Large Language Model Agents via Memory-Based Query Fragmentation and Fusion
Zixin Rao, Wentian Zhu, Chan Aristella Lu, Zhaorun Chen, Wei Niu, Le Guan, Bo Li, Zhen Xiang
摘要
Large language model (LLM) agents increasingly rely on long-term memory to support complex task execution, user personalization, and domain adaptation. Meanwhile, emerging access-control mechanisms for LLM agents are being explored to block policy-violating requests, aiming to prevent misuse and improve resource efficiency. In this paper, we reveal a novel attack surface arising from agents' memory operations: prohibited content triggering access control can be fragmented across interactions, stored in long-term memory in a benign-appearing form, and later reconstructed through memory retrieval, without appearing explicitly in the final user query. Specifically, we propose FragFuse, the first attack that enables unprivileged users to bypass agent access control by exploiting this temporal channel introduced by long-term memory. FragFuse operates in three stages: (1) identifying rejection-responsible fragments via black-box adaptive querying with fragment masking; (2) injecting these fragments into memory using marked carrier queries; and (3) retrieving and fusing the stored fragments through a follow-up attack query. While FragFuse can be instantiated manually for individual agents, we propose an optimization scheme that tunes fusion instructions and marker designs on surrogate models, enabling automated attack generation without violating the attacker's threat model assumptions. We evaluate FragFuse across four representative agent settings and task domains, covering three state-of-the-art agent access-control mechanisms. FragFuse achieves an average bypass success rate of 86.3% and an average end-to-end harmful task success rate of 41.1% across all settings, with only 4.4% average task success rate degradation compared to configurations without access control. Additionally, we show that alternative defenses, such as state-of-the-art prompt-injection detectors and perplexity detectors, cannot effectively address our attack.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 被引用 1,477 次
相关 Paper
- Unveiling Privacy Risks in LLM Agent MemoryBo Wang, Weiyi He, Shenglai Zeng, Zhen Xiang 等ACL 2025
- MemPoison: Bypassing Selective Memory Mechanisms to Plant Backdoors in LLM AgentsHongtao Wang, Se Yang, Yu Chen, Puzhuo LiuCCS 2026 · 被引用 5 次
- A-MemGuard: A Proactive Defense Framework For LLM-Based Agent MemoryQianshan Wei, Tengchao Yang, Yaochen Wang, Xinfeng Li 等ICML 2026
- DrunkAgent: Stealthy Memory Corruption in LLM-Powered Recommender AgentsShiyi Yang, Zhibo Hu, Xinshu Li, Chen Wang 等WWW 2026 · 被引用 6 次
- Memory Injection Attacks on LLM Agents via Query-Only InteractionShen Dong, Shaochen Xu, Pengfei He, Yige Li 等NeurIPS 2025 · 被引用 146 次
