Rescriber: Smaller-LLM-Powered User-Led Data Minimization for LLM-Based Chatbots
Jijie Zhou, Eryue Xu, Yaoyao Wu, Tianshi Li
Abstract
The proliferation of LLM-based conversational agents has resulted in excessive disclosure of identifiable or sensitive information.However, existing technologies fail to offer perceptible control or account for users' personal preferences about privacy-utility tradeoffs due to the lack of user involvement.To bridge this gap, we designed, built, and evaluated Rescriber, a browser extension that supports user-led data minimization in LLM-based conversational agents by helping users detect and sanitize personal information in their prompts.Our studies (N=Rescriber) showed that Rescriber helped users reduce unnecessary disclosure and addressed their privacy concerns.Users' subjective perceptions of the system powered by Llama3-8B were on par with that by GPT-4o.The comprehensiveness and consistency of the detection and sanitization emerge as essential factors that affect users' trust and perceived protection.Our findings confirm the viability of smaller-LLM-powered, userfacing, on-device privacy controls, presenting a promising approach to address the privacy and trust challenges of AI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 796ffe2e-6578-43a2-a12e-a7b0f3ba187cCited by top-tier papers9
- Operationalizing Data Minimization for Privacy-Preserving LLM PromptingJijie Zhou, Niloofar Mireshghallah, Tianshi LiICLR 2026 · 13 citations
- From Fragmentation to Integration: Exploring the Design Space of AI Agents for Human-as-the-Unit Privacy ManagementEryue Xu, Tianshi LiCHI 2026 · 3 citations
- Privasis: Synthesizing the Largest "Public" Private Dataset from ScratchHyunwoo Kim, Niloofar Mireshghallah, Michael Duan, Rui Xin et al.ICML 2026 · 3 citations
- PrivWeb: Unobtrusive and Content-aware Privacy Protection For Web AgentsShuning Zhang, Yutong Jiang, Rongjun Ma, Yuting Yang et al.CHI 2026 · 2 citations
- User Perceptions vs. Proxy LLM Judges: Privacy and Helpfulness in LLM Responses to Privacy-Sensitive ScenariosXiaoyuan Wu, Roshni Kaushik, Wenkai Li, Lujo Bauer et al.ACL 2026 · 2 citations
Builds on19
- Are Emergent Abilities of Large Language Models a Mirage?Rylan Schaeffer, Brando Miranda, Sanmi KoyejoNeurIPS 2023 · 796 citations
- Large Language Models Can Be Strong Differentially Private LearnersXuechen Li, Florian Tramèr, Percy Liang, Tatsunori HashimotoICLR 2022 · 502 citations
- Differentially Private Fine-tuning of Language ModelsDa Yu, Saurabh Naik, Arturs Backurs, Sivakanth Gopi et al.ICLR 2022 · 494 citations
- Deduplicating Training Data Mitigates Privacy Risks in Language ModelsNikhil Kandpal, Eric Wallace, Colin RaffelICML 2022 · 395 citations
- Beyond Memorization: Violating Privacy via Inference with Large Language ModelsRobin Staab, Mark Vero, Mislav Balunovic, Martin T. VechevICLR 2024 · 211 citations
Related papers
- Beyond PII: How Users Attempt to Estimate and Mitigate Implicit LLM InferenceSynthia Qia Wang, Sai Teja Peddinti, Nina Taft, Nick FeamsterCHI 2026 · 1 citation
- Malicious LLM-Based Conversational AI Makes Users Reveal Personal InformationXiao Zhan, Juan Carlos Carrillo, William Seymour, Jose SuchUSENIX Security 2025
- Helping Johnny Make Sense of Privacy Policies with LLMsVincent Freiberger, Arthur Fleig, Erik BuchmannCHI 2026 · 3 citations
- Unveiling Privacy Risks in LLM Agent MemoryBo Wang, Weiyi He, Shenglai Zeng, Zhen Xiang et al.ACL 2025
- Empowering Users in Digital Privacy Management through Interactive LLM-Based AgentsBolun Sun, Yifan Zhou, Haiyun JiangICLR 2025
