Privacy Collapse: Benign Fine-Tuning Can Break Contextual Privacy in Language Models
Anmol Goel, Cornelius Emde, Seong Joon Oh, Sangdoo Yun, Martin Gubri
摘要
We identify a novel phenomenon in language models: benign fine-tuning of frontier models can lead to privacy collapse. We find that diverse, subtle patterns in training data can degrade contextual privacy, including optimisation for helpfulness, exposure to user information, emotional and subjective dialogue, and debugging code printing internal variables, among others. Fine-tuned models lose their ability to reason about contextual privacy norms, share information inappropriately with tools, and violate memory boundaries across contexts. Privacy collapse is a "silent failure" because models maintain high performance on standard safety and utility benchmarks whilst exhibiting severe privacy vulnerabilities. Our experiments show evidence of privacy collapse across six models (closed and open weight), five fine-tuning datasets (real-world and controlled data), and two task categories (agentic and memory-based). Our mechanistic analysis reveals that privacy representations are uniquely fragile to fine-tuning, compared to task-relevant features which are preserved. Our results reveal a critical gap in current safety evaluations, in particular for the deployment of specialised agents. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow InstructionsFederico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger 等ICLR 2024 · 被引用 373 次
- In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space SteeringSheng Liu, Haotian Ye, Lei Xing, James Y. ZouICML 2024 · 被引用 244 次
- Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity TheoryNiloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov 等ICLR 2024 · 被引用 198 次
- AirGapAgent: Protecting Privacy-Conscious Conversational AgentsEugene Bagdasarian, Ren Yi, Sahra Ghalebikesabi, Peter Kairouz 等CCS 2024 · 被引用 5 次
- Leaky Thoughts: Large Reasoning Models Are Not Private ThinkersTommaso Green, Martin Gubri, Haritz Puerto, Sangdoo Yun 等EMNLP 2025 · 被引用 2 次
相关 Paper
- CIMemories: A Compositional Benchmark For Contextual Integrity In LLMsNiloofar Mireshghallah, Neal Mangaokar, Narine Kokhlikyan, Arman Zharmagambetov 等ICLR 2026 · 被引用 10 次
- The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy RisksXiaoyi Chen, Siyuan Tang, Rui Zhu, Shijun Yan 等CCS 2024 · 被引用 18 次
- The Heterogeneous Safety Impacts of Benign Multilingual Fine-TuningWill Hawkins, Kai Rawal, Jonathan Rystrøm, Stratis Tsirtsis 等ICML 2026
- Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained ModelsYuxin Wen, Leo Marchyok, Sanghyun Hong, Jonas Geiping 等NeurIPS 2024 · 被引用 39 次
- Large Language Models Can Be Contextual Privacy Protection LearnersYijia Xiao, Yiqiao Jin, Yushi Bai, Yue Wu 等EMNLP 2024 · 被引用 18 次
