Privacy Collapse: Benign Fine-Tuning Can Break Contextual Privacy in Language Models
Anmol Goel, Cornelius Emde, Seong Joon Oh, Sangdoo Yun, Martin Gubri
Abstract
We identify a novel phenomenon in language models: benign fine-tuning of frontier models can lead to privacy collapse. We find that diverse, subtle patterns in training data can degrade contextual privacy, including optimisation for helpfulness, exposure to user information, emotional and subjective dialogue, and debugging code printing internal variables, among others. Fine-tuned models lose their ability to reason about contextual privacy norms, share information inappropriately with tools, and violate memory boundaries across contexts. Privacy collapse is a "silent failure" because models maintain high performance on standard safety and utility benchmarks whilst exhibiting severe privacy vulnerabilities. Our experiments show evidence of privacy collapse across six models (closed and open weight), five fine-tuning datasets (real-world and controlled data), and two task categories (agentic and memory-based). Our mechanistic analysis reveals that privacy representations are uniquely fragile to fine-tuning, compared to task-relevant features which are preserved. Our results reveal a critical gap in current safety evaluations, in particular for the deployment of specialised agents. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 01f66a92-24c1-4c7c-b2a1-9a87077168bcBuilds on9
- Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow InstructionsFederico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger et al.ICLR 2024 · 373 citations
- In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space SteeringSheng Liu, Haotian Ye, Lei Xing, James Y. ZouICML 2024 · 244 citations
- Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity TheoryNiloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov et al.ICLR 2024 · 198 citations
- AirGapAgent: Protecting Privacy-Conscious Conversational AgentsEugene Bagdasarian, Ren Yi, Sahra Ghalebikesabi, Peter Kairouz et al.CCS 2024 · 5 citations
- Leaky Thoughts: Large Reasoning Models Are Not Private ThinkersTommaso Green, Martin Gubri, Haritz Puerto, Sangdoo Yun et al.EMNLP 2025 · 2 citations
Related papers
- CIMemories: A Compositional Benchmark For Contextual Integrity In LLMsNiloofar Mireshghallah, Neal Mangaokar, Narine Kokhlikyan, Arman Zharmagambetov et al.ICLR 2026 · 10 citations
- The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy RisksXiaoyi Chen, Siyuan Tang, Rui Zhu, Shijun Yan et al.CCS 2024 · 18 citations
- The Heterogeneous Safety Impacts of Benign Multilingual Fine-TuningWill Hawkins, Kai Rawal, Jonathan Rystrøm, Stratis Tsirtsis et al.ICML 2026
- Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained ModelsYuxin Wen, Leo Marchyok, Sanghyun Hong, Jonas Geiping et al.NeurIPS 2024 · 39 citations
- Large Language Models Can Be Contextual Privacy Protection LearnersYijia Xiao, Yiqiao Jin, Yushi Bai, Yue Wu et al.EMNLP 2024 · 18 citations
