User Perceptions vs. Proxy LLM Judges: Privacy and Helpfulness in LLM Responses to Privacy-Sensitive Scenarios
Xiaoyuan Wu, Roshni Kaushik, Wenkai Li, Lujo Bauer, Koichi Onoue
摘要
Large language models (LLMs) are rapidly being adopted for tasks like drafting emails, summarizing meetings, and answering health questions. In these settings, users may need to share private information (e.g., contact details, health records). To evaluate LLMs' ability to identify and redact such information, prior work introduced real-life, scenario-based benchmarks (e.g., ConfAIde, PrivacyLens) and found that LLMs can leak private information in complex scenarios. However, these evaluations relied on proxy LLMs to judge the helpfulness and privacy-preservation quality of LLM responses, rather than directly measuring users' perceptions. To understand how users perceive the helpfulness and privacy-preservation quality of LLM responses to privacy-sensitive scenarios, we conducted a user study (n = 94) using 90 PrivacyLens scenarios. We found that users had low agreement with each other when evaluating identical LLM responses. In contrast, five proxy LLMs reached high agreement, yet each proxy LLM had low correlation with users' evaluations. These results indicate that proxy LLMs cannot accurately estimate users' wide range of perceptions of utility and privacy in privacy-sensitive scenarios. We discuss the need for more user-centered studies to measure LLMs' ability to help users while preserving privacy, and for improving alignment between LLMs and users in estimating perceived privacy and utility.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity TheoryNiloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov 等ICLR 2024 · 被引用 198 次
- Quantifying Memorization Across Neural Language ModelsNicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee 等ICLR 2023 · 被引用 158 次
- The Feasibility of Dynamically Granted Permissions: Aligning Mobile Privacy with User PreferencesPrimal Wijesekera, Arjun Baokar, Lynn Tsai, Joel Reardon 等S&P 2017 · 被引用 156 次
- On the Robustness of Code Generation Techniques: An Empirical Study on GitHub CopilotAntonio Mastropaolo, Luca Pascarella, Emanuela Guglielmi, Matteo Ciniselli 等ICSE 2023 · 被引用 124 次
- Privacy Norms for Smart Home Personal AssistantsNoura Abdi, Xiao Zhan, Kopo M. Ramokapane, Jose M. SuchCHI 2021 · 被引用 107 次
相关 Paper
- Expectation Alignment of Language Models for Real-World User ExpectationsMiaomiao Li, Yang Wang, Bin Liang, Shudong Liu 等ICML 2026
- PII-Bench: Evaluating Query-Aware Privacy Protection SystemsHao Shen, Zhouhong Gu, Haokai Hong, Weili Han 等ACL 2026
- The Privacy Paradox of LLMs: User Perceptions and the Reality of PII LeakageShuai Cheng, Haitao Xu, Shu Meng, Shuai Hao 等CHI 2026
- A User-Centric Multi-Intent Benchmark for Evaluating Large Language ModelsJiayin Wang, Fengran Mo, Weizhi Ma, Peijie Sun 等EMNLP 2024 · 被引用 10 次
- Prevalence Overshadows Concerns? Understanding Chinese Users' Privacy Awareness and Expectations Towards LLM-Based Healthcare ConsultationZhihuang Liu, Ling Hu, Tongqing Zhou, Yonghao Tang 等S&P 2025
