Self-Refining Language Model Anonymizers via Adversarial Distillation
Kyuyoung Kim, Hyunjun Jeon, Jinwoo Shin
Abstract
Large language models (LLMs) are increasingly used in sensitive domains, where their ability to infer personal data from seemingly benign text introduces emerging privacy risks. While recent LLM-based anonymization methods help mitigate such risks, they often rely on proprietary models (e.g., GPT-4), raising concerns about cost and the potential exposure of sensitive data to untrusted external systems. To address this, we introduce SElf-refining Anonymization with Language model (SEAL), a novel distillation framework for training small language models (SLMs) to perform effective anonymization without relying on external models at inference time. SEAL leverages adversarial interactions between an LLM anonymizer and an inference model to collect trajectories of anonymized texts and inferred attributes, which are then used to distill anonymization and critique capabilities into SLMs through supervised fine-tuning and preference learning. The resulting models learn both to anonymize text and to evaluate their outputs, enabling iterative improvement of anonymization quality via self-refinement. Experiments on SynthPAI, a dataset of synthetic personal profiles and text comments, demonstrate that SLMs trained with SEAL achieve substantial improvements in anonymization capabilities. Notably, 8B models attain a privacy-utility trade-off comparable to that of the GPT-4 anonymizer and, with self-refinement, even surpass it in terms of privacy protection. These results highlight the effectiveness of our adversarial distillation framework for training SLMs as efficient anonymizers. * Equal contribution 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- RedacBench: Can AI Erase Your Secrets?Hyunjun Jeon, Kyuyoung Kim, Jinwoo ShinICLR 2026 · 2 citations
- Personalized Language Models via Privacy-Preserving Evolutionary Model MergingKyuyoung Kim, Jinwoo Shin, Jaehyung KimEMNLP 2025
Builds on11
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting et al.NeurIPS 2023 · 657 citations
- 8-bit Optimizers via Block-wise QuantizationTim Dettmers, Mike Lewis, Sam Shleifer, Luke ZettlemoyerICLR 2022 · 457 citations
- Beyond Memorization: Violating Privacy via Inference with Large Language ModelsRobin Staab, Mark Vero, Mislav Balunovic, Martin T. VechevICLR 2024 · 211 citations
Related papers
- Language Models are Advanced AnonymizersRobin Staab, Mark Vero, Mislav Balunovic, Martin T. VechevICLR 2025
- Self-Adapting Language ModelsAdam Zweiger, Jyothish Pari, Han Guo, Yoon Kim et al.NeurIPS 2025 · 78 citations
- Robust Utility-Preserving Text Anonymization Based on Large Language ModelsTianyu Yang, Xiaodan Zhu, Iryna GurevychACL 2025
- SEAL: Safety-enhanced Aligned LLM Fine-tuning via Bilevel Data SelectionHan Shen, Pin-Yu Chen, Payel Das, Tianyi ChenICLR 2025
- Self Iterative Label Refinement via Robust Unlabeled LearningHikaru Asano, Tadashi Kozuno, Yukino BabaNeurIPS 2025 · 1 citation
