Language Models are Advanced Anonymizers
Robin Staab, Mark Vero, Mislav Balunovic, Martin T. Vechev
Abstract
Recent privacy research on large language models (LLMs) has shown that they achieve near-human-level performance at inferring personal data from online texts. With ever-increasing model capabilities, existing text anonymization methods are currently lacking behind regulatory requirements and adversarial threats. In this work, we take two steps to bridge this gap: First, we present a new setting for evaluating anonymization in the face of adversarial LLM inferences, allowing for a natural measurement of anonymization performance while remedying some of the shortcomings of previous metrics. Then, within this setting, we develop a novel LLM-based adversarial anonymization framework leveraging the strong inferential capabilities of LLMs to inform our anonymization procedure. We conduct a comprehensive experimental evaluation of adversarial anonymization across 13 LLMs on real-world and synthetic online texts, comparing it against multiple baselines and industry-grade anonymizers. Our evaluation shows that adversarial anonymization outperforms current commercial anonymizers both in terms of the resulting utility and privacy. We support our findings with a human study (n=50) highlighting a strong and consistent human preference for LLM-anonymized texts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0a7521e6-538c-4485-b487-b040e89e11efCited by top-tier papers6
- AI and My Values: User Perceptions of LLMs' Ability to Extract, Embody, and Explain Human Values from Casual ConversationsBhada Yun, Renn Su, April Yi WangCHI 2026 · 7 citations
- Self-Refining Language Model Anonymizers via Adversarial DistillationKyuyoung Kim, Hyunjun Jeon, Jinwoo ShinNeurIPS 2025 · 7 citations
- Privasis: Synthesizing the Largest "Public" Private Dataset from ScratchHyunwoo Kim, Niloofar Mireshghallah, Michael Duan, Rui Xin et al.ICML 2026 · 3 citations
- Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMsDong Yan, Jian Liang, Ran He, Tieniu TanICLR 2026 · 3 citations
- RedacBench: Can AI Erase Your Secrets?Hyunjun Jeon, Kyuyoung Kim, Jinwoo ShinICLR 2026 · 2 citations
Builds on9
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting et al.NeurIPS 2023 · 657 citations
- Can Large Language Models Be an Alternative to Human Evaluations?David Cheng-Han Chiang, Hung-yi LeeACL 2023 · 254 citations
- Beyond Memorization: Violating Privacy via Inference with Large Language ModelsRobin Staab, Mark Vero, Mislav Balunovic, Martin T. VechevICLR 2024 · 211 citations
- Quantifying Memorization Across Neural Language ModelsNicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee et al.ICLR 2023 · 158 citations
Related papers
- Robust Utility-Preserving Text Anonymization Based on Large Language ModelsTianyu Yang, Xiaodan Zhu, Iryna GurevychACL 2025
- Private Attribute Inference from Images with Vision-Language ModelsBatuhan Tömekçe, Mark Vero, Robin Staab, Martin T. VechevNeurIPS 2024 · 54 citations
- De-Anonymization at Scale via Tournament-Style AttributionLirui Zhang, Huishuai ZhangACL 2026
- Reconstruction Attack-Resistant Inference Paradigm for LLM Cloud ServicesZipeng Ye, Wenjian Luo, Qi Zhou, Yubo TangAAAI 2026
- Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation GenerationAneta Zugecova, Dominik Macko, Ivan Srba, Róbert Móro et al.ACL 2025 · 18 citations
