Controlled Generation for Private Synthetic Text
Zihao Zhao, Anjalie Field
Abstract
Text anonymization is essential for responsibly developing and deploying AI in high-stakes domains such as healthcare, social services, and law. In this work, we propose a novel methodology for privacy-preserving synthetic text generation that leverages the principles of de-identification and the Hiding In Plain Sight (HIPS) theory. Our approach introduces entityaware control codes to guide controllable generation using either in-context learning (ICL) or prefix tuning. The ICL variant ensures privacy levels consistent with the underlying deidentification system, while the prefix tuning variant incorporates a custom masking strategy and loss function to support scalable, highquality generation. Experiments on legal and clinical datasets demonstrate that our method achieves a strong balance between privacy protection and utility, offering a practical and effective solution for synthetic text generation in sensitive domains. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 83916cbf-3c1e-4ebe-bcf7-9f69bfd80e21Builds on15
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun et al.NeurIPS 2021 · 606 citations
- Large Language Models Can Be Strong Differentially Private LearnersXuechen Li, Florian Tramèr, Percy Liang, Tatsunori HashimotoICLR 2022 · 502 citations
- Differentially Private Fine-tuning of Language ModelsDa Yu, Saurabh Naik, Arturs Backurs, Sivakanth Gopi et al.ICLR 2022 · 494 citations
Related papers
- Privacy Preserving In-Context-Learning Framework for Large Language ModelsBishnu Bhusal, Manoj Acharya, Ramneet Kaur, Colin Samplawski et al.AAAI 2026 · 1 citation
- Synthetic Text Generation with Differential Privacy: A Simple and Practical RecipeXiang Yue, Huseyin A. Inan, Xuechen Li, Girish Kumar et al.ACL 2023 · 24 citations
- Secret-Protected Evolution for Differentially Private Synthetic Text GenerationTianze Wang, Zhaoyu Chen, Jian Du, Yingtai Xiao et al.ICLR 2026 · 1 citation
- Disguise without Disruption: Utility-Preserving Face De-identificationZikui Cai, Zhongpai Gao, Benjamin Planche, Meng Zheng et al.AAAI 2024 · 28 citations
- PromptEHR: Conditional Electronic Healthcare Records Generation with Prompt LearningZifeng Wang, Jimeng SunEMNLP 2022 · 19 citations
