Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
Niloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov, Maarten Sap, Reza Shokri, Yejin Choi
Abstract
Existing efforts on quantifying privacy implications for large language models (LLMs) solely focus on measuring leakage of training data. In this work, we shed light on the often-overlooked interactive settings where an LLM receives information from multiple sources at inference time and generates an output to be shared with other entities, creating the potential of exposing sensitive input data in inappropriate contexts. In these scenarios, humans naturally uphold privacy by choosing whether or not to disclose information depending on the context. We ask the question "Can LLMs demonstrate an equivalent discernment and reasoning capability when considering privacy in context?" We propose CONFAIDE, a benchmark grounded in the theory of contextual integrity and designed to identify critical weaknesses in the privacy reasoning capabilities of instruction-tuned LLMs. CONFAIDE consists of four tiers, gradually increasing in complexity, with the final tier evaluating contextual privacy reasoning and theory of mind capabilities. Our experiments show that even commercial models such as GPT-4 and ChatGPT reveal private information in contexts that humans would not, 39% and 57% of the time, respectively, highlighting the urgent need for a new direction of privacy-preserving approaches as we demonstrate a larger underlying problem stemmed in the models' lack of reasoning capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cda0fa3c-2b97-4b77-b8dc-e20aca1b36ebCited by top-tier papers61
- Large Language Model Unlearning via Embedding-Corrupted PromptsChris Yuhao Liu, Yaxuan Wang, Jeffrey Flanigan, Yang LiuNeurIPS 2024 · 138 citations
- LLM-PBE: Assessing Data Privacy in Large Language ModelsQinbin Li, Junyuan Hong, Chulin Xie, Jeffrey Tan et al.VLDB 2024 · 66 citations
- Contextual Integrity in LLMs via Reasoning and Reinforcement LearningGuangchen Lan, Huseyin A. Inan, Sahar Abdelnabi, Janardhan Kulkarni et al.NeurIPS 2025 · 56 citations
- 'I'm Categorizing LLM as a Productivity Tool': Examining Ethics of LLM Use in HCI Research PracticesShivani Kapania, Ruiyi Wang, Toby Jia-Jun Li, Tianshi Li et al.CSCW 2025 · 31 citations
- SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak AttacksHongye Cao, Sijia Jing, Yanming Wang, Ziyue Peng et al.ICLR 2026 · 28 citations
Builds on16
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 1,715 citations
Related papers
- CIMemories: A Compositional Benchmark For Contextual Integrity In LLMsNiloofar Mireshghallah, Neal Mangaokar, Narine Kokhlikyan, Arman Zharmagambetov et al.ICLR 2026 · 10 citations
- PrivaCI-Bench: Evaluating Privacy with Contextual Integrity and Legal ComplianceHaoran Li, Wenbin Hu, Huihao Jing, Yulin Chen et al.ACL 2025
- ToMBench: Benchmarking Theory of Mind in Large Language ModelsZhuang Chen, Jincenzi Wu, Jinfeng Zhou, Bosi Wen et al.ACL 2024 · 6 citations
- User Perceptions vs. Proxy LLM Judges: Privacy and Helpfulness in LLM Responses to Privacy-Sensitive ScenariosXiaoyuan Wu, Roshni Kaushik, Wenkai Li, Lujo Bauer et al.ACL 2026 · 2 citations
- PrivLM-Bench: A Multi-level Privacy Evaluation Benchmark for Language ModelsHaoran Li, Dadi Guo, Donghao Li, Wei Fan et al.ACL 2024 · 9 citations
