Privasis: Synthesizing the Largest "Public" Private Dataset from Scratch
Hyunwoo Kim, Niloofar Mireshghallah, Michael Duan, Rui Xin, Stella Li, Jaehun Jung, David Acuna, Qi Pang, Hanshen Xiao, Edward Suh, Sewoong Oh, Yulia Tsvetkov
Abstract
Research involving privacy-sensitive data has always been constrained by data scarcity, standing in sharp contrast to other areas that have benefited from data scaling. This challenge is becoming increasingly urgent as modern AI agents-such as OpenClaw and Gemini Agent-are granted persistent access to highly sensitive personal information. To tackle this longstanding bottleneck and the rising risks, we present PRIVASIS (i.e., privacy oasis), the first million-scale fully synthetic dataset entirely built from scratch-an expansive reservoir of texts with rich and diverse private information-designed to broaden and accelerate research in areas where processing sensitive social data is inevitable. Compared to existing datasets, PRIVASIS, comprising 1.4 million records, offers orders-of-magnitude larger scale with quality, and far greater diversity across various document types, including medical history, legal documents, financial records, calendars, and text messages with a total of 55.1 million annotated attributes such as ethnicity, date of birth, workplace, etc. We leverage PRIVASIS to construct a parallel corpus for text sanitization with our pipeline that decomposes texts and applies targeted sanitization. Our compact sanitization models (≤4B) trained on this dataset outperform state-of-the-art large language models, such as GPT-5 and Qwen-3 235B. We plan to release data, models, and code to accelerate future research on privacy-sensitive domains and agents. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 55eb1927-9db5-4b4f-acf0-587601814f32Builds on18
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity TheoryNiloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov et al.ICLR 2024 · 198 citations
- Adapting Large Language Models via Reading ComprehensionDaixuan Cheng, Shaohan Huang, Furu WeiICLR 2024 · 146 citations
- Differentially Private Synthetic Data via Foundation Model APIs 2: TextChulin Xie, Zinan Lin, Arturs Backurs, Sivakanth Gopi et al.ICML 2024 · 71 citations
Related papers
- Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone AgentsZhixin Lin, Jungang Li, Shidong Pan, Yibo Shi et al.AAAI 2026 · 7 citations
- PrivSniffer: Graph-based Contextual Privacy Leakage Detection for User-Generated TextsHangyu Ye, Liyao Xiang, Naixuan Huang, Dongyue Yu et al.WWW 2026
- OmniSQL: Synthesizing High-quality Text-to-SQL Data at ScaleHaoyang Li, Shang Wu, Xiaokang Zhang, Xinmei Huang et al.VLDB 2025 · 90 citations
- How Private are Language Models in Abstractive Summarization?Anthony Hughes, Nikolaos Aletras, Ning MaEMNLP 2025
- PrivCode: When Code Generation Meets Differential PrivacyZheng Liu, Chen Gong, Terry Yue Zhuo, Kecen Li et al.NDSS 2026 · 5 citations
