Unveiling Privacy Risks in LLM Agent Memory
Bo Wang, Weiyi He, Shenglai Zeng, Zhen Xiang, Yue Xing, Jiliang Tang, Pengfei He
Abstract
Large Language Model (LLM) agents have become increasingly prevalent across various realworld applications. They enhance decisionmaking by storing private user-agent interactions in the memory module for demonstrations, introducing new privacy risks for LLM agents. In this work, we systematically investigate the vulnerability of LLM agents to our proposed Memory EXTRaction Attack (MEXTRA) under a black-box setting. To extract private information from memory, we propose an effective attacking prompt design and an automated prompt generation method based on different levels of knowledge about the LLM agent. Experiments on two representative agents demonstrate the effectiveness of MEXTRA. Moreover, we explore key factors influencing memory leakage from both the agent designer's and the attacker's perspectives. Our findings highlight the urgent need for effective memory safeguards in LLM agent design and deployment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6ee54f0b-303f-42bf-a57d-ab46368bc423Cited by top-tier papers13
- Privacy-Aware Decoding: Mitigating Privacy Leakage of Large Language Models in Retrieval-Augmented GenerationHaoran Wang, Xiongxiao Xu, Baixiang Huang, Kai ShuKDD 2026 · 13 citations
- Sponge Tool Attack: Stealthy Denial-of-Efficiency against Tool-Augmented Agentic ReasoningQi Li, Xinchao WangICML 2026 · 12 citations
- FragFuse: Bypassing Access Control of Large Language Model Agents via Memory-Based Query Fragmentation and FusionZixin Rao, Wentian Zhu, Chan Aristella Lu, Zhaorun Chen et al.USENIX Security 2026 · 5 citations
- From Fragmentation to Integration: Exploring the Design Space of AI Agents for Human-as-the-Unit Privacy ManagementEryue Xu, Tianshi LiCHI 2026 · 3 citations
- MemPot: Defend Against Memory Extraction Attack with Optimized HoneypotsYuhao Wang, Shengfang ZHAI, Guanghao Jin, Yinpeng Dong et al.ICML 2026 · 1 citation
Builds on14
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 1,477 citations
- GPT-4V(ision) is a Generalist Web Agent, if GroundedBoyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun et al.ICML 2024 · 496 citations
Related papers
- Exploiting the Shadows: Unveiling Privacy Leaks through Lower-Ranked Tokens in Large Language ModelsYuan Zhou, Zhuo Zhang, Xiangyu ZhangACL 2025 · 2 citations
- DrunkAgent: Stealthy Memory Corruption in LLM-Powered Recommender AgentsShiyi Yang, Zhibo Hu, Xinshu Li, Chen Wang et al.WWW 2026 · 6 citations
- Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMsXiang Zheng, YUTAO WU, Hanxun Huang, Yige Li et al.ICML 2026
- Memory Injection Attacks on LLM Agents via Query-Only InteractionShen Dong, Shaochen Xu, Pengfei He, Yige Li et al.NeurIPS 2025 · 146 citations
- Attractive Metadata Attack: Inducing LLM Agents to Invoke Malicious ToolsKanghua Mo, Li Hu, Yucheng Long, Zhihao LiNeurIPS 2025 · 37 citations
