The Privacy Paradox of LLMs: User Perceptions and the Reality of PII Leakage
Shuai Cheng, Haitao Xu, Shu Meng, Shuai Hao, Chuan Yue, Zhao Li
Abstract
Large language models (LLMs) are increasingly deployed, yet they introduce significant privacy risks by disclosing personally identifiable information (PII) during interactions. Although prior work has demonstrated the feasibility of extracting PII from LLMs, no comprehensive study has evaluated the actual extent of PII leakage across mainstream LLMs or investigated user perceptions, literacy, and behavioral responses to these risks. To address these gaps, we conduct a large-scale evaluation of PII leakage in popular LLMs, demonstrating that attackers can extract email addresses and phone numbers with high success rates. Through a mixed-methods study involving 20 interviews and 204 survey participants, we identify significant discrepancies between user concerns and behavior: despite strong concerns about PII leakage and limited understanding of training data provenance, users continue to use LLMs due to perceived utility, often exhibiting privacy cynicism. Based on these findings, we propose design implications for enhancing the privacyutility balance in future LLM deployments.
• Security and privacy → Human and societal aspects of security and privacy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9aa9f501-bcc2-4bae-aa82-c254457ea60cBuilds on28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos et al.USENIX Security 2019 · 1,386 citations
Related papers
- Effective PII Extraction from LLMs through Augmented Few-Shot LearningShuai Cheng, Shu Meng, Haitao Xu, Haoran Zhang et al.USENIX Security 2025
- "It's a Fair Game", or Is It? Examining How Users Navigate Disclosure Risks and Benefits When Using LLM-Based Conversational AgentsZhiping Zhang, Michelle Jia, Hao-Ping (Hank) Lee, Bingsheng Yao et al.CHI 2024 · 92 citations
- Teach LLMs to Phish: Stealing Private Information from Language ModelsAshwinee Panda, Christopher A. Choquette-Choo, Zhengming Zhang, Yaoqing Yang et al.ICLR 2024 · 41 citations
- User Perceptions vs. Proxy LLM Judges: Privacy and Helpfulness in LLM Responses to Privacy-Sensitive ScenariosXiaoyuan Wu, Roshni Kaushik, Wenkai Li, Lujo Bauer et al.ACL 2026 · 2 citations
- ProPILE: Probing Privacy Leakage in Large Language ModelsSiwon Kim, Sangdoo Yun, Hwaran Lee, Martin Gubri et al.NeurIPS 2023 · 229 citations
