Style Pooling: Automatic Text Style Obfuscation for Improved Classification Fairness
Fatemehsadat Mireshghallah, Taylor Berg-Kirkpatrick
摘要
Text style can reveal sensitive attributes of the author (e.g. race or age) to the reader, which can, in turn, lead to privacy violations and bias in both human and algorithmic decisions based on text. For example, the style of writing in job applications might reveal protected attributes of the candidate which could lead to bias in hiring decisions, regardless of whether hiring decisions are made algorithmically or by humans. We propose a VAE-based framework that obfuscates stylistic features of human-generated text through style transfer by automatically re-writing the text itself. Our framework operationalizes the notion of obfuscated style in a flexible way that enables two distinct notions of obfuscated style: (1) a minimal notion that effectively intersects the various styles seen in training, and (2) a maximal notion that seeks to obfuscate by adding stylistic features of all sensitive attributes to text, in effect, computing a union of styles. Our style-obfuscation framework can be used for multiple purposes, however, we demonstrate its effectiveness in improving the fairness of downstream classifiers. We also conduct a comprehensive study on style pooling's effect on fluency, semantic consistency, and attribute removal from text, in two and three domain style obfuscation. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Mix and Match: Learning-free Controllable Text Generationusing Energy Language ModelsFatemehsadat Mireshghallah, Kartik Goyal, Taylor Berg-KirkpatrickACL 2022 · 被引用 90 次
- Adversarial Authorship Attribution for DeobfuscationWanyue Zhai, Jonathan Rusert, Zubair Shafiq, Padmini SrinivasanACL 2022 · 被引用 7 次
它引用的顶会 Paper8
- A Probabilistic Formulation of Unsupervised Text Style TransferJunxian He, Xinyi Wang, Graham Neubig, Taylor Berg-KirkpatrickICLR 2020 · 被引用 136 次
- A4NT: Author Attribute Anonymity by Adversarial Training of Neural Machine TranslationRakshith Shetty, Bernt Schiele, Mario FritzUSENIX Security 2018 · 被引用 104 次
- PowerTransformer: Unsupervised Controllable Revision for Biased Language CorrectionXinyao Ma, Maarten Sap, Hannah Rashkin, Yejin ChoiEMNLP 2020 · 被引用 50 次
- Null It Out: Guarding Protected Attributes by Iterative Nullspace ProjectionShauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton 等ACL 2020 · 被引用 25 次
- Unsupervised Discovery of Implicit Gender BiasAnjalie Field, Yulia TsvetkovEMNLP 2020 · 被引用 5 次
相关 Paper
- A Causal Lens for Controllable Text GenerationZhiting Hu, Li Erran LiNeurIPS 2021 · 被引用 77 次
- Human-Guided Fair Classification for Natural Language ProcessingFlorian E. Dorner, Momchil Peychev, Nikola Konstantinov, Naman Goel 等ICLR 2023
- DP-VAE: Human-Readable Text Anonymization for Online Reviews with Differentially Private Variational AutoencodersBenjamin Weggenmann, Valentin Rublack, Michael Andrejczuk, Justus Mattern 等WWW 2022 · 被引用 55 次
- A Girl Has A Name: Detecting Authorship ObfuscationAsad Mahmood, Zubair Shafiq, Padmini SrinivasanACL 2020 · 被引用 1 次
- Revision in Continuous Space: Unsupervised Text Style Transfer without Adversarial LearningDayiheng Liu, Jie Fu, Yidan Zhang, Chris Pal 等AAAI 2020 · 被引用 53 次
