Anti-adversarial Learning: Desensitizing Prompts for Large Language Model
Xuan Li, Zhe Yin, Xiaodong Gu, Beijun Shen
Abstract
With the widespread use of LLMs, preserving privacy in user prompts has become crucial, as prompts risk exposing private and sensitive data to cloud LLMs. Conventional techniques like homomorphic encryption (HE), secure multi-party computation, and federated learning (FL) are not well-suited to this scenario due to the lack of control over user participation in remote model interactions. In this paper, we propose PromptObfus, a novel method for desensitizing LLM prompts. The core idea of PromptObfus is "anti-adversarial" learning. Unlike adversarial attacks that add imperceptible perturbations to mislead models, PromptObfus perturbs sensitive words to make them unrecognizable to humans while maintaining the model's original predictions. Specifically, PromptObfus frames prompt desensitization as a masked language modeling task, replacing privacy-sensitive terms with a [MASK] token. A desensitization model is utilized to generate candidate replacements for each masked position. These candidates are subsequently selected based on gradient feedback from a surrogate model, ensuring minimal disruption to task output. We demonstrate the effectiveness of our approach on three NLP tasks. Results show that PromptObfus effectively prevents privacy inference from remote LLMs while preserving task utility.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8742a644-e01a-4e10-96db-8727d21a75c1Builds on7
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu et al.ICLR 2024 · 817 citations
- Large Language Models are Human-Level Prompt EngineersYongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster et al.ICLR 2023 · 297 citations
- Iron: Private Inference on TransformersMeng Hao, Hongwei Li, Hanxiao Chen, Pengzhi Xing et al.NeurIPS 2022 · 209 citations
- Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity TheoryNiloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov et al.ICLR 2024 · 198 citations
- Automatic Prompt Optimization with "Gradient Descent" and Beam SearchReid Pryzant, Dan Iter, Jerry Li, Yin Tat Lee et al.EMNLP 2023 · 137 citations
Related papers
- ObfusLM: Privacy-preserving Language Model Service against Embedding Inversion AttacksYu Lin, Ruining Yang, Yunlong Mao, Qizhi Zhang et al.ACL 2025 · 2 citations
- Prompt Obfuscation for Large Language ModelsDavid Pape, Sina Mavali, Thorsten Eisenhofer, Lea SchönherrUSENIX Security 2025
- Prεεmpt: Sanitizing Sensitive Prompts for LLMsAmrita Roy Chowdhury, David Glukhov, Divyam Anshumaan, Prasad Chalasani et al.NDSS 2026 · 5 citations
- TextFusion: Privacy-Preserving Pre-trained Model Inference via Token FusionXin Zhou, Jinzhu Lu, Tao Gui, Ruotian Ma et al.EMNLP 2022 · 12 citations
- Your Inference Request Will Become a Black Box: Confidential Inference for Cloud-based Large Language ModelsChung-ju Huang, Huiqiang Zhao, Yuanpeng He, Lijian Li et al.ACL 2026
