Lune

AAAI2026Top-tier venue

Anti-adversarial Learning: Desensitizing Prompts for Large Language Model

Xuan Li, Zhe Yin, Xiaodong Gu, Beijun Shen

2026Year

Abstract

With the widespread use of LLMs, preserving privacy in user prompts has become crucial, as prompts risk exposing private and sensitive data to cloud LLMs. Conventional techniques like homomorphic encryption (HE), secure multi-party computation, and federated learning (FL) are not well-suited to this scenario due to the lack of control over user participation in remote model interactions. In this paper, we propose PromptObfus, a novel method for desensitizing LLM prompts. The core idea of PromptObfus is "anti-adversarial" learning. Unlike adversarial attacks that add imperceptible perturbations to mislead models, PromptObfus perturbs sensitive words to make them unrecognizable to humans while maintaining the model's original predictions. Specifically, PromptObfus frames prompt desensitization as a masked language modeling task, replacing privacy-sensitive terms with a [MASK] token. A desensitization model is utilized to generate candidate replacements for each masked position. These candidates are subsequently selected based on gradient feedback from a surrogate model, ensuring minimal disruption to task output. We demonstrate the effectiveness of our approach on three NLP tasks. Results show that PromptObfus effectively prevents privacy inference from remote LLMs while preserving task utility.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 8742a644-e01a-4e10-96db-8727d21a75c1

Builds on7

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines