SafetyMem: Adaptive Jailbreak Defense via Dual-Component Safety Memory
Hao Wang, Ziyi Ni, Huacan Wang, Pin Lyu, Lei Sha
Abstract
Current defenses for Large Language Models (LLMs) often suffer from a "memory gap": parameter-modifying methods are computationally rigid, while inference-time filters cannot retain or reuse defense knowledge across interactions. To address this, we propose Safet-yMem, a novel framework that secures LLMs through a dual-component safety memory system. SafetyMem consists of Semantic Safety Memory (SSM), which consolidates diverse jailbreak attempts into a structured knowledge base of attack patterns, and Episodic Safety Memory (ESM), which maintains an evolving set of procedural rules refined from historical detection failures. Unlike static defenses, Safe-tyMem allows the model to "remember" and adapt to emerging adversarial strategies without parameter retraining. To further enhance robustness, we introduce an adversarial memory expansion mechanism that proactively generates challenging variants to solidify these memories. Experiments on standard and stealthy jailbreak benchmarks show that SafetyMem substantially reduces attack success rates while preserving efficiency and interpretability, consistently outperforming state-of-the-art baselines across multiple LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e5fc8809-175f-43c3-9714-2329a23c24c5Builds on12
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 2,230 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
- Is ChatGPT a General-Purpose Natural Language Processing Task Solver?Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen et al.EMNLP 2023 · 449 citations
- GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via CipherYouliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang et al.ICLR 2024 · 441 citations
- Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow InstructionsFederico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger et al.ICLR 2024 · 373 citations
Related papers
- Towards Robust Multimodal Large Language Models Against Jailbreak AttacksZiyi Yin, Yuanpu Cao, Han Liu, Ting Wang et al.CVPR 2026 · 5 citations
- Stand on The Shoulders of Giants: Building JailExpert from Previous Attack ExperienceXi Wang, Songlei Jian, Shasha Li, Xiaopeng Li et al.EMNLP 2025 · 1 citation
- SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical MannerXunguang Wang, Daoyuan Wu, Zhenlan Ji, Zongjie Li et al.USENIX Security 2025
- MirrorShield: Towards Dynamic Adaptive Defense Against Jailbreaks via Entropy-Guided Mirror CraftingRui Pu, Chaozhuo Li, Rui Ha, Litian Zhang et al.AAAI 2026
- JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and ManipulationShenyi Zhang, Yuchen Zhai, Keyan Guo, Hongxin Hu et al.USENIX Security 2025
