DP-VAE: Human-Readable Text Anonymization for Online Reviews with Differentially Private Variational Autoencoders
Benjamin Weggenmann, Valentin Rublack, Michael Andrejczuk, Justus Mattern, Florian Kerschbaum
摘要
While vast amounts of personal data are shared daily on public online platforms and used by companies and analysts to gain valuable insights, privacy concerns are also on the rise: Modern authorship attribution techniques have proven effective at identifying individuals from their data, such as their writing style or behavior of picking and judging movies. It is hence crucial to develop data sanitization methods that allow sharing of users’ data while protecting their privacy and preserving quality and content of the original data. In this paper, we tackle anonymization of textual data and propose an end-to-end differentially private variational autoencoder architecture. Unlike previous approaches that achieve differential privacy on a per-word level through individual perturbations, our solution works at an abstract level by perturbing the latent vectors that provide a global summary of the input texts. Decoding an obfuscated latent vector thus not only allows our model to produce coherent, high-quality output text that is human-readable, but also results in strong anonymization due to the diversity of the produced data. We evaluate our approach on IMDb movie and Yelp business reviews, confirming its anonymization capabilities and preservation of the semantics and utility of the original sentences.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Synthetic Text Generation with Differential Privacy: A Simple and Practical RecipeXiang Yue, Huseyin A. Inan, Xuechen Li, Girish Kumar 等ACL 2023 · 被引用 24 次
- Differentially Private Language Models for Secure Data SharingJustus Mattern, Zhijing Jin, Benjamin Weggenmann, Bernhard Schölkopf 等EMNLP 2022 · 被引用 16 次
- FLAIM: AIM-based Synthetic Data Generation in the Federated SettingSamuel Maddock, Graham Cormode, Carsten MapleKDD 2024 · 被引用 5 次
- StyleRemix: Interpretable Authorship Obfuscation via Distillation and Perturbation of Style ElementsJillian Fisher, Skyler Hallinan, Ximing Lu, Mitchell L. Gordon 等EMNLP 2024 · 被引用 4 次
- CLOAK: Contrastive Guidance for Latent Diffusion-Based Data ObfuscationXin Yang, Omid ArdakanianUbiComp 2026
它引用的顶会 Paper3
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- A4NT: Author Attribute Anonymity by Adversarial Training of Neural Machine TranslationRakshith Shetty, Bernt Schiele, Mario FritzUSENIX Security 2018 · 被引用 104 次
相关 Paper
- Style Pooling: Automatic Text Style Obfuscation for Improved Classification FairnessFatemehsadat Mireshghallah, Taylor Berg-KirkpatrickEMNLP 2021 · 被引用 6 次
- Explainable Disentangled Representation Learning for Generalizable Authorship Attribution in the Era of Generative AIHieu Man, Van-Cuong Pham, Nghia Trung Ngo, Franck Dernoncourt 等ACL 2026
- A Neural Approach to Spatio-Temporal Data Release with User-Level Differential PrivacyRitesh Ahuja, Sepanta Zeighami, Gabriel Ghinita, Cyrus ShahabiSIGMOD 2023 · 被引用 14 次
- De-Anonymization at Scale via Tournament-Style AttributionLirui Zhang, Huishuai ZhangACL 2026
- Unsupervised Opinion Summarization as Copycat-Review GenerationArthur Brazinskas, Mirella Lapata, Ivan TitovACL 2020 · 被引用 14 次
