Prism: Private Relational Data Synthesis with Language Models
Guohui Guan, Chang Ge
Abstract
Organizations increasingly rely on relational data containing sensitive personal information for downstream tasks. A common approach is to release synthetic data as a privacy-preserving substitute, typically by learning a generative model and sampling from it. Differential privacy (DP) provides a formal guarantee of data privacy, and existing data synthesis systems aim to balance its inherent trade-off between data privacy and task utility. The advent of pretrained large language models (LLMs) has reshaped this trade-off landscape: on the utility side, LLMs possess stronger representational capabilities and encode extensive prior knowledge that does not consume the privacy budget, leading to improved task utility; on the other hand, enforcing differential privacy for LLM-based relational data synthesis introduces new technical challenges. We present P rism, an end-to-end framework for DP relational data synthesis using LLMs. P rism privatizes knowledge transfer from an ensemble of teacher LLMs to a student generator via DP output aggregation. We introduce three key innovations in the aggregation mechanism to address LLM-specific challenges, including 1) token pruning with tries to minimize teacher queries and hence the privacy cost; 2) a data-dependent DP mechanism that adaptively scales noise for efficient budget use; and 3) probabilistic sampling with threshold filtering to reduce bias and preserve diversity. Extensive empirical results across twelve experiments show that Prism achieves substantially higher predictive utility than eight state-of-the-art methods under the same privacy budget.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get bdb26984-ecfa-41fb-85fa-fe0f952da79dRelated papers
- Contrastive Private Data Synthesis via Weighted Multi-PLM FusionTianyuan Zou, Yang Liu, Peng Li, Yufei Xiong et al.ICML 2025
- Privacy Preserving In-Context-Learning Framework for Large Language ModelsBishnu Bhusal, Manoj Acharya, Ramneet Kaur, Colin Samplawski et al.AAAI 2026 · 1 citation
- Differentially Private Preference Data Synthesis for Large Language Model AlignmentFengyu Gao, Jing YangICML 2026
- PEARL: Differentially Private and Entropy-Aware Regulated Language GenerationSeongho Joo, Hyukhun Koh, Kyomin JungICML 2026
- PrivCode: When Code Generation Meets Differential PrivacyZheng Liu, Chen Gong, Terry Yue Zhuo, Kecen Li et al.NDSS 2026 · 5 citations
