Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents
Shuai Shao, Qihan Ren, Dongrui Liu, Chen Qian, Boyi Wei, Dadi Guo, Yang JingYi, Xinhao Song, Linfeng Zhang, Weinan Zhang, Jing Shao
Abstract
Advances in Large Language Models (LLMs) have enabled a new class of selfevolving agents that autonomously improve through environmental interaction, demonstrating strong capabilities. However, self-evolution also introduces novel risks overlooked by current safety research. In this work, we study case where an agent's self-evolution deviates in unintended ways, leading to undesirable or even harmful outcomes. We refer to this as Misevolution. We evaluate misevolution along four key evolutionary pathways: model, memory, tool, and workflow. Our empirical findings reveal that misevolution is a widespread risk, affecting agents built even on top-tier LLMs (e.g., Gemini-2.5-Pro). Different emergent risks are observed, such as degradation of safety alignment after memory accumulation, or unintended introduction of vulnerabilities in tool creation and reuse. To our knowledge, this is the first study to systematically conceptualize misevolution and provide empirical evidence of its occurrence, highlighting an urgent need for new safety paradigms for self-evolving agents. Finally, we discuss potential mitigation strategies to inspire further research on building safer and more trustworthy selfevolving agents. Our code is available here. INTRODUCTION Large Language Model (LLM) agents are increasingly deployed in real-world applications, such as software development and automated research (Hong et al., 2024; OpenAI, 2025b). Recently, a new frontier focuses on agents that can evolve on their own, known as self-evolving agents (Zhou et al., 2025b; Zhang et al., 2025a; Gao et al., 2025; Fang et al., 2025) . Different from their static counterparts, these agents improve themselves via active and continuous interaction with the environment. The evolutionary process of these agents primarily spans four dimensions, each corresponding to a core component of the agent system: model, memory, tool, and workflow. By leveraging feedback from tasks, the agent may optimize the parameters of the underlying language model (Sun et al., 2025b), accumulate experience into memory (Zhou et al., 2025a), create and master new tools (Qiu et al., 2025) , or adjust the execution workflow (Zhang et al., 2025b). The impressive performance of self-evolving agents on challenging tasks has drawn wide interest in the community. However, self-evolution also introduces novel risks that are overlooked by current safety research. In this study, we investigate the case in which an agent's self-evolution deviates in unintended ways, leading to undesirable or even harmful outcomes. We refer to this as Misevolution, and highlight four core characteristics that distinguish it from established safety concerns: 1. Temporal emergence. During self-evolution, some components of the agent are dynamically changing, and risks can emerge over time. This contrasts with research on jailbreaking or misalignment that evaluates a "static snapshot" of an LLM (Chao et al., 2024; Li et al., 2023) .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 12bee63f-0f0f-46c4-b180-9ce9b86d331fCited by top-tier papers6
- Panacea: Mitigating Harmful Fine-tuning for Large Language Models via Post-fine-tuning PerturbationYibo Wang, Tiansheng Huang, Li Shen, Huanjin Yao et al.NeurIPS 2025 · 22 citations
- Antibody: Strengthening Defense Against Harmful Fine-Tuning for Large Language Models via Attenuating Harmful Gradient InfluenceQuoc Minh Nguyen, Trung Le, Jing Wu, Anh Tuan Bui et al.ICLR 2026 · 10 citations
- Are Your Agents Upward Deceivers?Dadi Guo, Qingyu Liu, Dongrui Liu, Qihan Ren et al.ICML 2026 · 5 citations
- Surgery: Mitigating Harmful Fine-Tuning for Large Language Models via Attention SinkGuozhi Liu, Weiwei Lin, Tiansheng Huang, Ruichao Mo et al.ICML 2026 · 5 citations
- MonoScale: Scaling Multi-Agent System with Monotonic ImprovementShuai Shao, Yixiang Liu, Bingwei Lu, Weinan ZhangICML 2026
Builds on27
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen et al.ICLR 2024 · 1,104 citations
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun et al.ICLR 2024 · 716 citations
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge BasesZhaorun Chen, Zhen Xiang, Chaowei Xiao, Dawn Song et al.NeurIPS 2024 · 539 citations
Related papers
- AGrail: A Lifelong Agent Guardrail with Effective and Adaptive Safety DetectionWeidi Luo, Shenghong Dai, Xiaogeng Liu, Suman Banerjee et al.ACL 2025 · 41 citations
- AgentBreeder: Mitigating the AI Safety Risks of Multi-Agent Scaffolds via Self-ImprovementJ. Rosser, Jakob N. FoersterNeurIPS 2025 · 12 citations
- Gödel Agent: A Self-Referential Agent Framework for Recursively Self-ImprovementXunjian Yin, Xinyi Wang, Liangming Pan, Li Lin et al.ACL 2025 · 7 citations
- Identifying the Risks of LM Agents with an LM-Emulated SandboxYangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis et al.ICLR 2024 · 292 citations
- Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMsXiang Zheng, YUTAO WU, Hanxun Huang, Yige Li et al.ICML 2026
