A4NT: Author Attribute Anonymity by Adversarial Training of Neural Machine Translation
Rakshith Shetty, Bernt Schiele, Mario Fritz
摘要
Text-based analysis methods enable an adversary to reveal privacy relevant author attributes such as gender, age and can identify the text's author. Such methods can compromise the privacy of an anonymous author even when the author tries to remove privacy sensitive content. In this paper, we propose an automatic method, called Adversarial Author Attribute Anonymity Neural Translation (A 4 NT), to combat such text-based adversaries. Unlike prior works on obfuscation, we propose a system that is fully automatic and learns to perform obfuscation entirely from data. This allows us to easily apply the A 4 NT system to obfuscate different author attributes. We propose a sequence-to-sequence language model, inspired by machine translation, and an adversarial training framework to design a system which learns to transform the input text to obfuscate author attributes without paired data. We also propose and evaluate techniques to impose constraints on our A 4 NT to preserve the semantics of the input text. A 4 NT learns to make minimal changes to the input to successfully fool author attribute classifiers, while preserving the meaning of the input text. Our experiments on two datasets and three settings show that the proposed method is effective in fooling the author attribute classifiers and thus improves the anonymity of authors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Privacy Risks of General-Purpose Language ModelsXudong Pan, Mi Zhang, Shouling Ji, Min YangS&P 2020 · 被引用 291 次
- Adversarial Watermarking Transformer: Towards Tracing Text Provenance with Data HidingSahar Abdelnabi, Mario FritzS&P 2021 · 被引用 210 次
- Information Leakage in Embedding ModelsCongzheng Song, Ananth RaghunathanCCS 2020 · 被引用 200 次
- Realtime Robust Malicious Traffic Detection via Frequency Domain AnalysisChuanpu Fu, Qi Li, Meng Shen, Ke XuCCS 2021 · 被引用 194 次
- DP-VAE: Human-Readable Text Anonymization for Online Reviews with Differentially Private Variational AutoencodersBenjamin Weggenmann, Valentin Rublack, Michael Andrejczuk, Justus Mattern 等WWW 2022 · 被引用 55 次
相关 Paper
- Adversarial Authorship Attribution for DeobfuscationWanyue Zhai, Jonathan Rusert, Zubair Shafiq, Padmini SrinivasanACL 2022 · 被引用 7 次
- A Girl Has A Name: Detecting Authorship ObfuscationAsad Mahmood, Zubair Shafiq, Padmini SrinivasanACL 2020 · 被引用 1 次
- ALISON: Fast and Effective Stylometric Authorship ObfuscationEric Xing, Saranya Venkatraman, Thai Le, Dongwon LeeAAAI 2024 · 被引用 7 次
- Style Pooling: Automatic Text Style Obfuscation for Improved Classification FairnessFatemehsadat Mireshghallah, Taylor Berg-KirkpatrickEMNLP 2021 · 被引用 6 次
- Language Models are Advanced AnonymizersRobin Staab, Mark Vero, Mislav Balunovic, Martin T. VechevICLR 2025
