A-MemGuard: A Proactive Defense Framework For LLM-Based Agent Memory
Qianshan Wei, Tengchao Yang, Yaochen Wang, Xinfeng Li, Lijun Li, Zhenfei Yin, Yi Zhan, Thorsten Holz, Zhiqiang Lin, XiaoFeng Wang
Abstract
Large Language Model (LLM) agents use memory to learn from past interactions, enabling autonomous planning and decision-making in complex environments. However, this reliance on memory introduces a critical security risk: an adversary can inject seemingly harmless records into an agent's memory to manipulate its future behavior. This vulnerability is characterized by two core aspects: First, the malicious effect of injected records is only activated within a specific context, making them hard to detect when individual memory entries are audited in isolation. Second, once triggered, the manipulation can initiate a self-reinforcing error cycle: the corrupted outcome is stored as precedent, which not only amplifies the initial error but also progressively lowers the threshold for similar attacks in the future. To address these challenges, we introduce A-MemGuard (Agent-Memory Guard), the first proactive defense framework for LLM agent memory. The core idea of our work is the insight that memory itself must become both self-checking and self-correcting. Without modifying the agent's core architecture, A-MemGuard combines two mechanisms: (1) consensus-based validation, which detects anomalies by comparing reasoning paths derived from multiple related memories and (2) a dual-memory structure, where detected failures are distilled into "lessons" stored separately and consulted before future actions, breaking error cycles and enabling adaptation. Comprehensive evaluations on multiple benchmarks show that A-MemGuard effectively cuts attack success rates by over 95% while incurring a minimal utility cost. This work shifts LLM memory security from static filtering to a proactive, experience-driven model where defenses strengthen over time. Our code is available in https://github.com/TangciuYueng/AMemGuard
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 74fde67f-9fa7-4b7b-8ec6-76be73fae842Cited by top-tier papers2
- Your Agent May Misevolve: Emergent Risks in Self-evolving LLM AgentsShuai Shao, Qihan Ren, Dongrui Liu, Chen Qian et al.ICLR 2026 · 60 citations
- Visual Inception: Compromising Long-term Planning in Agentic Recommenders via Multimodal Memory PoisoningJiachen QianACL 2026 · 3 citations
Builds on17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao et al.ICCV 2023 · 685 citations
Related papers
- Safeguarding LLM Agents against Long-Horizon Threats via Shadow MemoryYuhui Wang, Tanqiu Jiang, Jiacheng Liang, Charles Fleming et al.CCS 2026
- Memory Injection Attacks on LLM Agents via Query-Only InteractionShen Dong, Shaochen Xu, Pengfei He, Yige Li et al.NeurIPS 2025 · 146 citations
- MemPoison: Bypassing Selective Memory Mechanisms to Plant Backdoors in LLM AgentsHongtao Wang, Se Yang, Yu Chen, Puzhuo LiuCCS 2026 · 5 citations
- SafetyMem: Adaptive Jailbreak Defense via Dual-Component Safety MemoryHao Wang, Ziyi Ni, Huacan Wang, Pin Lyu et al.ACL 2026
- FragFuse: Bypassing Access Control of Large Language Model Agents via Memory-Based Query Fragmentation and FusionZixin Rao, Wentian Zhu, Chan Aristella Lu, Zhaorun Chen et al.USENIX Security 2026 · 5 citations
