ACL2026

PerMemSafe: Benchmarking Implicit Personalized Safety of Long Horizon Self-Evolving Agents

Hengyu An, Minxi Li, Naen Xu, Chunyi Zhou, Xiaogang Xu, Tianyu Du, Jinbao Li, Shouling Ji

Abstract

Self-evolving agents achieve personalization by accumulating user-specific memories over long horizons. This capability, however, introduces novel safety risks, as responses that are generally safe may become harmful in userspecific contexts. Such safety-relevant contexts often emerge implicitly and evolve over time during long-horizon conversations, rendering traditional context-independent safety evaluations insufficient. To address this, we formally define Implicit Personalized Safety and present PERMEMSAFE, the first benchmark for evaluating implicit personalized safety of self-evolving agents in long-horizon interactions. Empirical results reveal significant limitations of existing self-evolving agents, with even the strongest achieving only around 50% safety rate, highlighting systematic failures in reasoning about personalized safety risks. To mitigate this, we propose SENTINELMEM, an active risk-aware memory framework that explicitly models personalized risk inference and memory evolution. Experiments show that SENTINELMEM improves implicit personalized safety by 23.8% over prior memory frameworks while maintaining helpfulness in long horizon interactions. 1