Secret-Protected Evolution for Differentially Private Synthetic Text Generation
Tianze Wang, Zhaoyu Chen, Jian Du, Yingtai Xiao, Linjun Zhang, Qiang Yan
Abstract
Text data has become extremely valuable on large language models (LLMs) and even lead to general artificial intelligence (AGI). A lot of high-quality text in the real world is private and cannot be freely used due to privacy concerns. Therefore, differentially private (DP) synthetic text generation has been proposed, aiming to produce high-utility synthetic data while protecting sensitive information. However, existing DP synthetic text generation imposes uniform guarantees that often overprotect non-sensitive content, resulting in substantial utility loss and computational overhead. Therefore, we propose Secret-Protected Evolution (SecPE), a novel framework that extends private evolution with secret-aware protection. Theoretically, we show that SecPE satisfies (p, r)-secret protection, constituting a relaxation of Gaussian DP that enables tighter utility-privacy trade-offs, while also substantially reducing computational complexity relative to baseline methods. Empirically, across the OpenReview, PubMed, and Yelp benchmarks, SecPE consistently achieves lower Fréchet Inception Distance (FID) and higher downstream task accuracy than GDP-based Aug-PE baselines, while requiring less noise to attain the same level of protection. Our results highlight that secret-aware guarantees can unlock more practical and effective privacy-preserving synthetic text generation. * indicates equal contributions. This work was done when Tianze Wang was an intern at TikTok.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2e59e0d5-9b29-4479-82f9-caad41592489Builds on12
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos et al.USENIX Security 2019 · 1,386 citations
- Reconstructing Training Data with Informed AdversariesBorja Balle, Giovanni Cherubin, Jamie HayesS&P 2022 · 214 citations
- Learning with User-Level PrivacyDaniel Levy, Ziteng Sun, Kareem Amin, Satyen Kale et al.NeurIPS 2021 · 113 citations
- Bounding training data reconstruction in DP-SGDJamie Hayes, Borja Balle, Saeed MahloujifarNeurIPS 2023 · 73 citations
Related papers
- Differentially Private Synthetic Data via Foundation Model APIs 2: TextChulin Xie, Zinan Lin, Arturs Backurs, Sivakanth Gopi et al.ICML 2024 · 71 citations
- Private Evolution ConvergesTomás González Lara, Giulia Fanti, Aaditya RamdasNeurIPS 2025 · 3 citations
- Differentially Private Synthetic Data via Foundation Model APIs 1: ImagesZinan Lin, Sivakanth Gopi, Janardhan Kulkarni, Harsha Nori et al.ICLR 2024 · 63 citations
- SeqPATE: Differentially Private Text Generation via Knowledge DistillationZhiliang Tian, Yingxiu Zhao, Ziyue Huang, Yu-Xiang Wang et al.NeurIPS 2022 · 29 citations
- Privacy Preserving In-Context-Learning Framework for Large Language ModelsBishnu Bhusal, Manoj Acharya, Ramneet Kaur, Colin Samplawski et al.AAAI 2026 · 1 citation
