Investigating Novice Researchers' Perceptions of Research Privacy Within LLM-Assisted Workflows
Shuning Zhang, Changxi Wen, Eve He, Ying Ma, Robert Xiao, Xin Yi, Hewu Li
Abstract
Large Language Model (LLMs)-assisted scholarly workflows introduce critical privacy and intellectual property risks. As a uniquely vulnerable cohort driven by publication pressure and a lack of institutional support, novice researchers rely heavily on public LLMs, compelling them to navigate high-stakes privacy-publication tradeoffs. To investigate these concerns, we conducted semi-structured interviews with 44 researchers across diverse disciplines. Our findings reveal that the fear of idea leakage paradoxically accelerates, rather than deters, reliance on LLMs, as researchers utilize them to expedite publication. They also held misconceptions that their ideas lacked the unique value to attract targeted attacks, and that their inputs would be safely diluted within massive datasets, preventing reconstruction. From interviews, we identified five types of mitigations including input fragmentation and adversarial probing, though we found that participants largely perceived these measures as ineffective. We outline implications including implementing institution-level sandboxed isolation, scenario-based privacy pedagogy, and verifiable data-deletion audits for transparency.
• Security and privacy → Usability in security and privacy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6568d0f8-3030-45df-b93c-67614084e944Builds on35
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Formalizing and Benchmarking Prompt Injection Attacks and DefensesYupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia et al.USENIX Security 2024 · 308 citations
- "It's a Fair Game", or Is It? Examining How Users Navigate Disclosure Risks and Benefits When Using LLM-Based Conversational AgentsZhiping Zhang, Michelle Jia, Hao-Ping (Hank) Lee, Bingsheng Yao et al.CHI 2024 · 92 citations
- DE-COP: Detecting Copyrighted Content in Language Models Training DataAndré V. Duarte, Xuandong Zhao, Arlindo L. Oliveira, Lei LiICML 2024 · 81 citations
- Privacy Champions in Software Teams: Understanding Their Motivations, Strategies, and ChallengesMohammad Tahaei, Alisa Frik, Kami VanieaCHI 2021 · 75 citations
Related papers
- The Privacy Paradox of LLMs: User Perceptions and the Reality of PII LeakageShuai Cheng, Haitao Xu, Shu Meng, Shuai Hao et al.CHI 2026
- Evaluating Large Language Models on Academic Literature Understanding and Review: An Empirical Study among Early-stage ScholarsJiyao Wang, Haolong Hu, Zuyuan Wang, Song Yan et al.CHI 2024 · 21 citations
- LLM-PBE: Assessing Data Privacy in Large Language ModelsQinbin Li, Junyuan Hong, Chulin Xie, Jeffrey Tan et al.VLDB 2024 · 66 citations
- Creative Writers' Attitudes on Writing as Training Data for Large Language ModelsKaty Ilonka Gero, Meera A. Desai, Carly Schnitzler, Nayun Eom et al.CHI 2025 · 15 citations
- Privy: Envisioning and Mitigating Privacy Risks for Consumer-facing AI Product ConceptsHao-Ping (Hank) Lee, Yu-Ju Yang, Matthew Bilik, Isadora Krsek et al.CHI 2026 · 1 citation
