KnowledgeSG: Privacy-Preserving Synthetic Text Generation with Knowledge Distillation from Server
Wenhao Wang, Xiaoyu Liang, Rui Ye, Jingyi Chai, Siheng Chen, Yanfeng Wang
摘要
The success of large language models (LLMs) facilitate many parties to fine-tune LLMs on their own private data. However, this practice raises privacy concerns due to the memorization of LLMs. Existing solutions, such as utilizing synthetic data for substitution, struggle to simultaneously improve performance and preserve privacy. They either rely on a local model for generation, resulting in a performance decline, or take advantage of APIs, directly exposing the data to API servers. To address this issue, we propose KnowledgeSG, a novel client-server framework which enhances synthetic data quality and improves model performance while ensuring privacy. We achieve this by learning local knowledge from the private data with differential privacy (DP) and distilling professional knowledge from the server. Additionally, inspired by federated learning, we transmit models rather than data between the client and server to prevent privacy leakage. Extensive experiments in medical and financial domains demonstrate the effectiveness of Knowl-edgeSG. Our code is now publicly available at https://github.com/wwh0411/KnowledgeSG .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- MCP-Flow: Facilitating LLM Agents to Master Real-World, Diverse and Scaling MCP ToolsWenhao Wang, Peizhi Niu, Zhao Xu, Zhaoyu Chen 等ACL 2026 · 被引用 8 次
- MCP-Persona: Benchmarking LLM Agents on Real-World Personal Applications via Environment SimulationWenhao Wang, Peizhi Niu, Gongyi Zou, Xiyuan Yang 等ICML 2026 · 被引用 1 次
- RewardDS: Privacy-Preserving Fine-Tuning for Large Language Models via Reward Driven Data SynthesisJianwei Wang, Chengming Shi, Junyao Yang, Haoran Li 等EMNLP 2025
- FedMABench: Benchmarking Mobile GUI Agents on Decentralized Heterogeneous User DataWenhao Wang, Zijie Yu, Rui Ye, Jianqing Zhang 等EMNLP 2025
它引用的顶会 Paper20
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun 等ICLR 2024 · 被引用 945 次
相关 Paper
- Privacy Preserving In-Context-Learning Framework for Large Language ModelsBishnu Bhusal, Manoj Acharya, Ramneet Kaur, Colin Samplawski 等AAAI 2026 · 被引用 1 次
- Model-based Large Language Model Customization as ServiceZhaomin Wu, Jizhou Guo, Junyi Hou, Bingsheng He 等EMNLP 2025
- PPC-GPT: Federated Task-Specific Compression of Large Language Models via Pruning and Chain-of-Thought DistillationTao Fan, Guoqiang Ma, Yuanfeng Song, Lixin Fan 等EMNLP 2025 · 被引用 1 次
- Flick: Empowering Federated Learning with Commonsense KnowledgeRan Zhu, Mingkun Yang, Shiqiang Wang, Jie Yang 等NeurIPS 2025
- Enhancing Small Medical Learners with Privacy-preserving Contextual PromptingXinlu Zhang, Shiyang Li, Xianjun Yang, Chenxin Tian 等ICLR 2024 · 被引用 13 次
