Investigating Novice Researchers' Perceptions of Research Privacy Within LLM-Assisted Workflows
Shuning Zhang, Changxi Wen, Eve He, Ying Ma, Robert Xiao, Xin Yi, Hewu Li
摘要
Large Language Model (LLMs)-assisted scholarly workflows introduce critical privacy and intellectual property risks. As a uniquely vulnerable cohort driven by publication pressure and a lack of institutional support, novice researchers rely heavily on public LLMs, compelling them to navigate high-stakes privacy-publication tradeoffs. To investigate these concerns, we conducted semi-structured interviews with 44 researchers across diverse disciplines. Our findings reveal that the fear of idea leakage paradoxically accelerates, rather than deters, reliance on LLMs, as researchers utilize them to expedite publication. They also held misconceptions that their ideas lacked the unique value to attract targeted attacks, and that their inputs would be safely diluted within massive datasets, preventing reconstruction. From interviews, we identified five types of mitigations including input fragmentation and adversarial probing, though we found that participants largely perceived these measures as ineffective. We outline implications including implementing institution-level sandboxed isolation, scenario-based privacy pedagogy, and verifiable data-deletion audits for transparency.
• Security and privacy → Usability in security and privacy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper35
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Formalizing and Benchmarking Prompt Injection Attacks and DefensesYupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia 等USENIX Security 2024 · 被引用 308 次
- "It's a Fair Game", or Is It? Examining How Users Navigate Disclosure Risks and Benefits When Using LLM-Based Conversational AgentsZhiping Zhang, Michelle Jia, Hao-Ping (Hank) Lee, Bingsheng Yao 等CHI 2024 · 被引用 92 次
- DE-COP: Detecting Copyrighted Content in Language Models Training DataAndré V. Duarte, Xuandong Zhao, Arlindo L. Oliveira, Lei LiICML 2024 · 被引用 81 次
- Privacy Champions in Software Teams: Understanding Their Motivations, Strategies, and ChallengesMohammad Tahaei, Alisa Frik, Kami VanieaCHI 2021 · 被引用 75 次
相关 Paper
- The Privacy Paradox of LLMs: User Perceptions and the Reality of PII LeakageShuai Cheng, Haitao Xu, Shu Meng, Shuai Hao 等CHI 2026
- Evaluating Large Language Models on Academic Literature Understanding and Review: An Empirical Study among Early-stage ScholarsJiyao Wang, Haolong Hu, Zuyuan Wang, Song Yan 等CHI 2024 · 被引用 21 次
- LLM-PBE: Assessing Data Privacy in Large Language ModelsQinbin Li, Junyuan Hong, Chulin Xie, Jeffrey Tan 等VLDB 2024 · 被引用 66 次
- Creative Writers' Attitudes on Writing as Training Data for Large Language ModelsKaty Ilonka Gero, Meera A. Desai, Carly Schnitzler, Nayun Eom 等CHI 2025 · 被引用 15 次
- Privy: Envisioning and Mitigating Privacy Risks for Consumer-facing AI Product ConceptsHao-Ping (Hank) Lee, Yu-Ju Yang, Matthew Bilik, Isadora Krsek 等CHI 2026 · 被引用 1 次
