EPSVec: Efficient and Private Synthetic Data Generation via Dataset Vectors
Mohammadamin Banayeeanzade, Qingchuan Yang, Deqing Fu, Spencer Hong, Erin Babinsky, Alfy Samuel, Anoop Kumar, Robin Jia, Sai Praneeth Reddy Karimireddy
Abstract
High-quality data is essential for modern machine learning, yet many valuable corpora are sensitive and cannot be freely shared. Synthetic data offers a practical substitute for downstream development, and large language models (LLMs) have emerged as powerful engines for generating it. However, existing private text generation methods are severely inefficient: they are data-intensive, computationally slow, and often require large private corpora or batch sizes to achieve usable quality. We introduce EPSVec, a differentially-private lightweight alternative that steers LLM generation using dataset vectors -directions in activation space that capture the distributional gap between private data and public priors. EPSVec extracts and sanitizes steering vectors just once and then performs standard decoding. This decouples the privacy budget from generation, enabling arbitrarily many synthetic samples without additional privacy cost and yielding strong fidelity even in low-data regimes. Furthermore, we enhance our method by utilizing pretrained (base) models and introducing fixed-shot prompting to boost generation diversity and fidelity. Our experiments demonstrate that EPSVec outperforms existing baselines in distributional alignment and downstream utility, particularly in low-data regimes, while significantly reducing computational overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on13
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun et al.NeurIPS 2021 · 606 citations
- Differentially Private Fine-tuning of Language ModelsDa Yu, Saurabh Naik, Arturs Backurs, Sivakanth Gopi et al.ICLR 2022 · 494 citations
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 461 citations
- Flocks of Stochastic Parrots: Differentially Private Prompt Learning for Large Language ModelsHaonan Duan, Adam Dziedzic, Nicolas Papernot, Franziska BoenischNeurIPS 2023 · 116 citations
- DP-OPT: Make Large Language Model Your Privacy-Preserving Prompt EngineerJunyuan Hong, Jiachen T. Wang, Chenhui Zhang, Zhangheng Li et al.ICLR 2024 · 70 citations
Related papers
- PrE-Text: Training Language Models on Private Federated Data in the Age of LLMsCharlie Hou, Akshat Shrivastava, Hongyuan Zhan, Rylan Conway et al.ICML 2024 · 30 citations
- Synthesizing Privacy-Preserving Text Data via Finetuning without Finetuning Billion-Scale LLMsBowen Tan, Zheng Xu, Eric P. Xing, Zhiting Hu et al.ICML 2025
- Differentially Private Steering for Large Language Model AlignmentAnmol Goel, Yaxi Hu, Iryna Gurevych, Amartya SanyalICLR 2025
- Differentially Private Synthetic Data via Foundation Model APIs 2: TextChulin Xie, Zinan Lin, Arturs Backurs, Sivakanth Gopi et al.ICML 2024 · 71 citations
- Synthetic Text Generation with Differential Privacy: A Simple and Practical RecipeXiang Yue, Huseyin A. Inan, Xuechen Li, Girish Kumar et al.ACL 2023 · 24 citations
