Privacy-Preserving Domain Adaptation of Semantic Parsers
Fatemehsadat Mireshghallah, Yu Su, Tatsunori Hashimoto, Jason Eisner, Richard Shin
Abstract
Task-oriented dialogue systems often assist users with personal or confidential matters. For this reason, the developers of such a system are generally prohibited from observing actual usage. So how can they know where the system is failing and needs more training data or new functionality? In this work, we study ways in which realistic user utterances can be generated synthetically, to help increase the linguistic and functional coverage of the system, without compromising the privacy of actual users. To this end, we propose a two-stage Differentially Private (DP) generation method which first generates latent semantic parses, and then generates utterances based on the parses. Our proposed approach improves MAUVE by 2.5× and parse tree function type overlap by 1.3× relative to current approaches for private synthetic data generation, improving both on fluency and semantic coverage. We further validate our approach on a realistic domain adaptation task of adding new functionality from private user data to a semantic parser, and show overall gains of 8.5% points in accuracy with the new feature.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Privacy-Preserving In-Context Learning with Differentially Private Few-Shot GenerationXinyu Tang, Richard Shin, Huseyin A. Inan, Andre Manoel et al.ICLR 2024 · 111 citations
- Synthetic Text Generation with Differential Privacy: A Simple and Practical RecipeXiang Yue, Huseyin A. Inan, Xuechen Li, Girish Kumar et al.ACL 2023 · 24 citations
- CopyBench: Measuring Literal and Non-Literal Reproduction of Copyright-Protected Text in Language Model GenerationTong Chen, Akari Asai, Niloofar Mireshghallah, Sewon Min et al.EMNLP 2024 · 4 citations
- Privacy Preserving In-Context-Learning Framework for Large Language ModelsBishnu Bhusal, Manoj Acharya, Ramneet Kaur, Colin Samplawski et al.AAAI 2026 · 1 citation
- Data-adaptive Differentially Private Prompt Synthesis for In-Context LearningFengyu Gao, Ruida Zhou, Tianhao Wang, Cong Shen et al.ICLR 2025
Builds on23
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos et al.USENIX Security 2019 · 1,386 citations
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun et al.NeurIPS 2021 · 606 citations
- Large Language Models Can Be Strong Differentially Private LearnersXuechen Li, Florian Tramèr, Percy Liang, Tatsunori HashimotoICLR 2022 · 502 citations
Related papers
- ACTG-ARL: Differentially Private Conditional Text Generation with RL-Boosted ControlYuzheng Hu, Ryan McKenna, Da Yu, Shanshan Wu et al.ICML 2026 · 1 citation
- MALA: Cross-Domain Dialogue Generation with Action LearningXinting Huang, Jianzhong Qi, Yu Sun, Rui ZhangAAAI 2020 · 19 citations
- Counterfactual Data Augmentation via Perspective Transition for Open-Domain DialoguesJiao Ou, Jinchao Zhang, Yang Feng, Jie ZhouEMNLP 2022 · 9 citations
- Differentially Private Preference Data Synthesis for Large Language Model AlignmentFengyu Gao, Jing YangICML 2026
- Differentially Private Language Models for Secure Data SharingJustus Mattern, Zhijing Jin, Benjamin Weggenmann, Bernhard Schölkopf et al.EMNLP 2022 · 16 citations
