When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents
Jiahe Guo, Xiangran Guo, Yulin Hu, Zimo Long, Xingyu Sui, Xuda Zhi, Yongbo Huang, Hao He, Weixiang Zhao, Yanyan Zhao, Bing Qin
Abstract
Long-term memory enables large language model (LLM) agents to support personalized and sustained interactions. However, most work on personalized agents prioritizes utility and user experience, treating memory as a neutral component and largely overlooking its safety implications. In this paper, we reveal intent legitimation, a previously underexplored safety failure in personalized agents, where benign personal memories bias intent inference and cause models to legitimize inherently harmful queries. To study this phenomenon, we introduce PS-Bench, a benchmark designed to identify and quantify intent legitimation in personalized interactions. Across multiple memory-augmented agent frameworks and base LLMs, personalization increases attack success rates by 15.8%--243.7% relative to stateless baselines. We further provide mechanistic evidence for intent legitimation from internal representations space, and propose a lightweight detection-reflection method that effectively reduces safety degradation. Overall, our work provides the first systematic exploration and evaluation of intent legitimation as a safety failure mode that naturally arises from benign, real-world personalization, highlighting the importance of assessing safety under long-term personal context. Our code is available at: https://github.com/MuyuenLP/PS-Bench. WARNING: This paper may contain harmful content.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c9715af8-5224-4ea4-9121-12304805699dBuilds on12
- A-Mem: Agentic Memory for LLM AgentsWujiang Xu, Zujie Liang, Kai Mei, Hang Gao et al.NeurIPS 2025 · 1,138 citations
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge BasesZhaorun Chen, Zhen Xiang, Chaowei Xiao, Dawn Song et al.NeurIPS 2024 · 539 citations
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 461 citations
- MemoryBank: Enhancing Large Language Models with Long-Term MemoryWanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye et al.AAAI 2024 · 394 citations
- Uncovering Safety Risks of Large Language Models through Concept Activation VectorZhihao Xu, Ruixuan Huang, Changyu Chen, Xiting WangNeurIPS 2024 · 83 citations
Related papers
- PersistBench: When Should Long-Term Memories Be Forgotten by LLMs?Sidharth Pulipaka, Oliver Chen, Manas Sharma, Taaha Saleem Bajwa et al.ICML 2026
- AMA-Bench: Evaluating Long-Horizon Memory for Agentic ApplicationsYujie Zhao, Boqin Yuan, Junbo Huang, Haocheng Yuan et al.ICML 2026 · 40 citations
- Unveiling Privacy Risks in LLM Agent MemoryBo Wang, Weiyi He, Shenglai Zeng, Zhen Xiang et al.ACL 2025
- Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based AgentsHanrong Zhang, Jingyuan Huang, Kai Mei, Yifei Yao et al.ICLR 2025
- Safeguarding LLM Agents against Long-Horizon Threats via Shadow MemoryYuhui Wang, Tanqiu Jiang, Jiacheng Liang, Charles Fleming et al.CCS 2026
