What Users Ask, Policies Miss: Unveiling the Gap Between Community-Expressed Privacy Concerns and LLM Provider Policies
Zhihuang Liu, Zhen Huang, Ling Hu, Yifan Yang, Zhiping Cai
摘要
Large Language Models(LLMs) process millions of conversations containing sensitive information daily, yet whether their privacy policies adequately address users' publicly expressed concerns remains unexplored. This paper presents the first large-scale, user-centered audit of privacy policy adequacy in LLM services. We systematically extract privacy concerns from Reddit communities of five major LLM providers and assess whether their latest privacy policies address these concerns. Our semi-automated pipeline analyzes 1,531 threads to extract 4,994 authentic privacy concerns, which we organize into a 20-topic taxonomy spanning four thematic groups. We then identify 3,137 policy gap instances and classify them into six categories: four policy coverage gaps (detail vague, AI feature unaddressed, vulnerable group neglected, and jurisdiction unclear) and two user perception gaps (explicit distrust and awareness deficit). Our analysis reveals that coverage gaps and perception gaps contribute nearly equally (50.2% vs. 49.8%), with AI-specific feature gaps (36.6%) and user awareness deficits (40.7%) together comprising the vast majority (77.3%) of all identified gaps. Notably, we also surface user-reported evidence suggesting potential discrepancies between stated policies and observed system behavior, highlighting the need for verifiable privacy guarantees. These findings demonstrate that improving LLM privacy requires dual-pronged interventions addressing both inadequate policy disclosures and user comprehension barriers, offering actionable insights for relevant stakeholders.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper38
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- NExT-GPT: Any-to-Any Multimodal LLMShengqiong Wu, Hao Fei, Leigang Qu, Wei Ji 等ICML 2024 · 被引用 786 次
- Polisis: Automated Analysis and Presentation of Privacy Policies Using Deep LearningHamza Harkous, Kassem Fawaz, Rémi Lebret, Florian Schaub 等USENIX Security 2018 · 被引用 400 次
- How Well Do My Results Generalize? Comparing Security and Privacy Survey Results from MTurk, Web, and Telephone SamplesElissa M. Redmiles, Sean Kross, Michelle L. MazurekS&P 2019 · 被引用 222 次
- PolicyLint: Investigating Internal Privacy Policy Contradictions on Google PlayBenjamin Andow, Samin Yaseer Mahmud, Wenyu Wang, Justin Whitaker 等USENIX Security 2019 · 被引用 185 次
相关 Paper
- Privacy Control in Conversational LLM Platforms: A Walkthrough StudyZhuoyang Li, Yanlai Wu, Yao Li, Xinning Gui 等CHI 2026 · 被引用 1 次
- Beyond Memorization: Violating Privacy via Inference with Large Language ModelsRobin Staab, Mark Vero, Mislav Balunovic, Martin T. VechevICLR 2024 · 被引用 211 次
- Prevalence Overshadows Concerns? Understanding Chinese Users' Privacy Awareness and Expectations Towards LLM-Based Healthcare ConsultationZhihuang Liu, Ling Hu, Tongqing Zhou, Yonghao Tang 等S&P 2025
- The Privacy Paradox of LLMs: User Perceptions and the Reality of PII LeakageShuai Cheng, Haitao Xu, Shu Meng, Shuai Hao 等CHI 2026
- Exploring User Security and Privacy Attitudes and Concerns Toward the Use of General-Purpose LLM Chatbots for Mental HealthJabari Kwesi, Jiaxun Cao, Riya Manchanda, Pardis Emami NaeiniUSENIX Security 2025
