Reducing Privacy Risks in Online Self-Disclosures with Language Models
Yao Dou, Isadora Krsek, Tarek Naous, Anubha Kabra, Sauvik Das, Alan Ritter, Wei Xu
Abstract
Self-disclosure, while being common and rewarding in social media interaction, also poses privacy risks. In this paper, we take the initiative to protect the user-side privacy associated with online self-disclosure through detection and abstraction. We develop a taxonomy of 19 self-disclosure categories and curate a large corpus consisting of 4.8K annotated disclosure spans. We then fine-tune a language model for detection, achieving over 65% partial span F 1 . We further conduct an HCI user study, with 82% of participants viewing the model positively, highlighting its real-world applicability. Motivated by the user feedback, we introduce the task of self-disclosure abstraction, which is rephrasing disclosures into less specific terms while preserving their utility, e.g., "Im 16F" to "I'm a teenage girl". We explore various fine-tuning strategies, and our best model can generate diverse abstractions that moderately reduce privacy risks while maintaining high utility according to human evaluation. To help users in deciding which disclosures to abstract, we present a task of rating their importance for context understanding. Our fine-tuned model achieves 80% accuracy, on par with GPT-3.5. Given safety and privacy considerations, we will only release our corpus and models to researchers who agree to the ethical guidelines outlined in our Ethics Statement. 1 * Appearance : "I am 6 '2". * Pet : " I have two musk turtles " * Occupation : "I 'm a motorcycle tourer ( by profession ) " , student should be categorized as Education . * Education : " I got accepted to UCLA " * Finance : any financial situations , not necessarily exact amounts .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e75f5dc2-16b3-433a-91cd-23ebf4211d2aCited by top-tier papers16
- Large-scale online deanonymization with LLMsSimon Lermen, Daniel Paleka, Joshua Swanson, Michael Aerni et al.USENIX Security 2026 · 20 citations
- Rescriber: Smaller-LLM-Powered User-Led Data Minimization for LLM-Based ChatbotsJijie Zhou, Eryue Xu, Yaoyao Wu, Tianshi LiCHI 2025 · 15 citations
- Operationalizing Data Minimization for Privacy-Preserving LLM PromptingJijie Zhou, Niloofar Mireshghallah, Tianshi LiICLR 2026 · 13 citations
- PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context OptimizationYidan Wang, Yanan Cao, Yubing Ren, Fang Fang et al.ACL 2025 · 12 citations
- Self-Refining Language Model Anonymizers via Adversarial DistillationKyuyoung Kim, Hyunjun Jeon, Jinwoo ShinNeurIPS 2025 · 7 citations
Builds on18
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Large Language Models Can Be Strong Differentially Private LearnersXuechen Li, Florian Tramèr, Percy Liang, Tatsunori HashimotoICLR 2022 · 502 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
- ProPILE: Probing Privacy Leakage in Large Language ModelsSiwon Kim, Sangdoo Yun, Hwaran Lee, Martin Gubri et al.NeurIPS 2023 · 229 citations
Related papers
- Measuring, Modeling, and Helping People Account for Privacy Risks in Online Self-Disclosures with AIIsadora Krsek, Anubha Kabra, Yao Dou, Tarek Naous et al.CSCW 2025 · 6 citations
- What Users Ask, Policies Miss: Unveiling the Gap Between Community-Expressed Privacy Concerns and LLM Provider PoliciesZhihuang Liu, Zhen Huang, Ling Hu, Yifan Yang et al.USENIX Security 2026
- SoK: A Privacy Framework for Security Research Using Social Media DataKyle Beadle, Kieron Ivy Turk, Aliai Eusebi, Mindy Tran et al.S&P 2025
- Supporting Informed Self-Disclosure: Design Recommendations for Presenting AI-Estimates of Privacy Risks to UsersIsadora Krsek, Meryl Ye, Wei Xu, Alan Ritter et al.CHI 2026 · 1 citation
- PrivSniffer: Graph-based Contextual Privacy Leakage Detection for User-Generated TextsHangyu Ye, Liyao Xiang, Naixuan Huang, Dongyue Yu et al.WWW 2026
