Privasis: Synthesizing the Largest "Public" Private Dataset from Scratch
Hyunwoo Kim, Niloofar Mireshghallah, Michael Duan, Rui Xin, Stella Li, Jaehun Jung, David Acuna, Qi Pang, Hanshen Xiao, Edward Suh, Sewoong Oh, Yulia Tsvetkov
摘要
Research involving privacy-sensitive data has always been constrained by data scarcity, standing in sharp contrast to other areas that have benefited from data scaling. This challenge is becoming increasingly urgent as modern AI agents-such as OpenClaw and Gemini Agent-are granted persistent access to highly sensitive personal information. To tackle this longstanding bottleneck and the rising risks, we present PRIVASIS (i.e., privacy oasis), the first million-scale fully synthetic dataset entirely built from scratch-an expansive reservoir of texts with rich and diverse private information-designed to broaden and accelerate research in areas where processing sensitive social data is inevitable. Compared to existing datasets, PRIVASIS, comprising 1.4 million records, offers orders-of-magnitude larger scale with quality, and far greater diversity across various document types, including medical history, legal documents, financial records, calendars, and text messages with a total of 55.1 million annotated attributes such as ethnicity, date of birth, workplace, etc. We leverage PRIVASIS to construct a parallel corpus for text sanitization with our pipeline that decomposes texts and applies targeted sanitization. Our compact sanitization models (≤4B) trained on this dataset outperform state-of-the-art large language models, such as GPT-5 and Qwen-3 235B. We plan to release data, models, and code to accelerate future research on privacy-sensitive domains and agents. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
- Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity TheoryNiloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov 等ICLR 2024 · 被引用 198 次
- Adapting Large Language Models via Reading ComprehensionDaixuan Cheng, Shaohan Huang, Furu WeiICLR 2024 · 被引用 146 次
- Differentially Private Synthetic Data via Foundation Model APIs 2: TextChulin Xie, Zinan Lin, Arturs Backurs, Sivakanth Gopi 等ICML 2024 · 被引用 71 次
相关 Paper
- Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone AgentsZhixin Lin, Jungang Li, Shidong Pan, Yibo Shi 等AAAI 2026 · 被引用 7 次
- PrivSniffer: Graph-based Contextual Privacy Leakage Detection for User-Generated TextsHangyu Ye, Liyao Xiang, Naixuan Huang, Dongyue Yu 等WWW 2026
- OmniSQL: Synthesizing High-quality Text-to-SQL Data at ScaleHaoyang Li, Shang Wu, Xiaokang Zhang, Xinmei Huang 等VLDB 2025 · 被引用 90 次
- How Private are Language Models in Abstractive Summarization?Anthony Hughes, Nikolaos Aletras, Ning MaEMNLP 2025
- PrivCode: When Code Generation Meets Differential PrivacyZheng Liu, Chen Gong, Terry Yue Zhuo, Kecen Li 等NDSS 2026 · 被引用 5 次
