Searching for Privacy Risks in LLM Agents via Simulation
Yanzhe Zhang, Diyi Yang
摘要
The widespread deployment of LLM-based agents is likely to introduce a critical privacy threat: malicious agents that proactively engage others in multi-turn interactions to extract sensitive information. However, the evolving nature of such dynamic dialogues makes it challenging to anticipate emerging vulnerabilities and design effective defenses. To tackle this problem, we present a search-based framework that alternates between improving attack and defense strategies through the simulation of privacy-critical agent interactions. Specifically, we employ LLMs as optimizers to analyze simulation trajectories and iteratively propose new agent instructions. To explore the strategy space more efficiently, we further utilize parallel search with multiple threads and cross-thread propagation. Through this process, we find that attack strategies escalate from direct requests to sophisticated tactics, such as impersonation and consent forgery, while defenses evolve from simple rule-based constraints to robust identity-verification state machines. The discovered attacks and defenses generalize across diverse scenarios and backbone models, providing useful insights for developing privacy-aware agents 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- PrivAct: Internalizing Contextual Privacy Preservation via Multi-Agent Preference TrainingYuhan Cheng, Hancheng Ye, Hai Li, Jingwei Sun 等ICML 2026 · 被引用 3 次
- Do Vision-Language Models Respect Contextual Integrity in Location Disclosure?Ruixin Yang, Ethan Mendes, Arthur Wang, James Hays 等ICLR 2026 · 被引用 1 次
- PriMobiBench: Characterizing Visual Privacy Leakage in VLM-Driven Mobile GUI AgentsQihang Cen, Tianshuo Cong, Da Song, Xinlei He 等CCS 2026
- Contextualized Privacy Defense for LLM AgentsYule Wen, Yanzhe Zhang, Jianxun Lian, Xiaoyuan Yi 等ICML 2026
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach 等ICLR 2022 · 被引用 1,976 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu 等ICLR 2024 · 被引用 817 次
相关 Paper
- Speculative Safety Honeypot: Toward Proactive Defense Against Multi-turn Agent AttacksZezhong WANG, Xueyang Tang, RUI LIAN, Yang Lou 等ICML 2026
- Unveiling Privacy Risks in LLM Agent MemoryBo Wang, Weiyi He, Shenglai Zeng, Zhen Xiang 等ACL 2025
- MetaMancer: Practical Tool Selection Attacks against LLM AgentsWenbin Zhai, Yihao Liu, Dong Wang, Hao Chen 等CCS 2026
- When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent SystemsHaowen Xu, Xue Tan, Lei Ma, Zhihao Zhang 等ICML 2026
- Stop Fixating on Prompts: Reasoning Hijacking and Constraint Tightening for Red-Teaming LLM AgentsYanxu Mao, Peipei Liu, Tiehan Cui, Congying Liu 等ACL 2026
