Searching for Privacy Risks in LLM Agents via Simulation
Yanzhe Zhang, Diyi Yang
Abstract
The widespread deployment of LLM-based agents is likely to introduce a critical privacy threat: malicious agents that proactively engage others in multi-turn interactions to extract sensitive information. However, the evolving nature of such dynamic dialogues makes it challenging to anticipate emerging vulnerabilities and design effective defenses. To tackle this problem, we present a search-based framework that alternates between improving attack and defense strategies through the simulation of privacy-critical agent interactions. Specifically, we employ LLMs as optimizers to analyze simulation trajectories and iteratively propose new agent instructions. To explore the strategy space more efficiently, we further utilize parallel search with multiple threads and cross-thread propagation. Through this process, we find that attack strategies escalate from direct requests to sophisticated tactics, such as impersonation and consent forgery, while defenses evolve from simple rule-based constraints to robust identity-verification state machines. The discovered attacks and defenses generalize across diverse scenarios and backbone models, providing useful insights for developing privacy-aware agents 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 569c174c-0151-41fc-8c9f-fd7d40bb598dCited by top-tier papers4
- PrivAct: Internalizing Contextual Privacy Preservation via Multi-Agent Preference TrainingYuhan Cheng, Hancheng Ye, Hai Li, Jingwei Sun et al.ICML 2026 · 3 citations
- Do Vision-Language Models Respect Contextual Integrity in Location Disclosure?Ruixin Yang, Ethan Mendes, Arthur Wang, James Hays et al.ICLR 2026 · 1 citation
- PriMobiBench: Characterizing Visual Privacy Leakage in VLM-Driven Mobile GUI AgentsQihang Cen, Tianshuo Cong, Da Song, Xinlei He et al.CCS 2026
- Contextualized Privacy Defense for LLM AgentsYule Wen, Yanzhe Zhang, Jianxun Lian, Xiaoyuan Yi et al.ICML 2026
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu et al.ICLR 2024 · 817 citations
Related papers
- Speculative Safety Honeypot: Toward Proactive Defense Against Multi-turn Agent AttacksZezhong WANG, Xueyang Tang, RUI LIAN, Yang Lou et al.ICML 2026
- Unveiling Privacy Risks in LLM Agent MemoryBo Wang, Weiyi He, Shenglai Zeng, Zhen Xiang et al.ACL 2025
- MetaMancer: Practical Tool Selection Attacks against LLM AgentsWenbin Zhai, Yihao Liu, Dong Wang, Hao Chen et al.CCS 2026
- When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent SystemsHaowen Xu, Xue Tan, Lei Ma, Zhihao Zhang et al.ICML 2026
- Stop Fixating on Prompts: Reasoning Hijacking and Constraint Tightening for Red-Teaming LLM AgentsYanxu Mao, Peipei Liu, Tiehan Cui, Congying Liu et al.ACL 2026
