Style Pooling: Automatic Text Style Obfuscation for Improved Classification Fairness
Fatemehsadat Mireshghallah, Taylor Berg-Kirkpatrick
Abstract
Text style can reveal sensitive attributes of the author (e.g. race or age) to the reader, which can, in turn, lead to privacy violations and bias in both human and algorithmic decisions based on text. For example, the style of writing in job applications might reveal protected attributes of the candidate which could lead to bias in hiring decisions, regardless of whether hiring decisions are made algorithmically or by humans. We propose a VAE-based framework that obfuscates stylistic features of human-generated text through style transfer by automatically re-writing the text itself. Our framework operationalizes the notion of obfuscated style in a flexible way that enables two distinct notions of obfuscated style: (1) a minimal notion that effectively intersects the various styles seen in training, and (2) a maximal notion that seeks to obfuscate by adding stylistic features of all sensitive attributes to text, in effect, computing a union of styles. Our style-obfuscation framework can be used for multiple purposes, however, we demonstrate its effectiveness in improving the fairness of downstream classifiers. We also conduct a comprehensive study on style pooling's effect on fluency, semantic consistency, and attribute removal from text, in two and three domain style obfuscation. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 921789df-eb00-4068-b73b-b455f55ab8dfCited by top-tier papers2
- Mix and Match: Learning-free Controllable Text Generationusing Energy Language ModelsFatemehsadat Mireshghallah, Kartik Goyal, Taylor Berg-KirkpatrickACL 2022 · 90 citations
- Adversarial Authorship Attribution for DeobfuscationWanyue Zhai, Jonathan Rusert, Zubair Shafiq, Padmini SrinivasanACL 2022 · 7 citations
Builds on8
- A Probabilistic Formulation of Unsupervised Text Style TransferJunxian He, Xinyi Wang, Graham Neubig, Taylor Berg-KirkpatrickICLR 2020 · 136 citations
- A4NT: Author Attribute Anonymity by Adversarial Training of Neural Machine TranslationRakshith Shetty, Bernt Schiele, Mario FritzUSENIX Security 2018 · 104 citations
- PowerTransformer: Unsupervised Controllable Revision for Biased Language CorrectionXinyao Ma, Maarten Sap, Hannah Rashkin, Yejin ChoiEMNLP 2020 · 50 citations
- Null It Out: Guarding Protected Attributes by Iterative Nullspace ProjectionShauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton et al.ACL 2020 · 25 citations
- Unsupervised Discovery of Implicit Gender BiasAnjalie Field, Yulia TsvetkovEMNLP 2020 · 5 citations
Related papers
- A Causal Lens for Controllable Text GenerationZhiting Hu, Li Erran LiNeurIPS 2021 · 77 citations
- Human-Guided Fair Classification for Natural Language ProcessingFlorian E. Dorner, Momchil Peychev, Nikola Konstantinov, Naman Goel et al.ICLR 2023
- DP-VAE: Human-Readable Text Anonymization for Online Reviews with Differentially Private Variational AutoencodersBenjamin Weggenmann, Valentin Rublack, Michael Andrejczuk, Justus Mattern et al.WWW 2022 · 55 citations
- A Girl Has A Name: Detecting Authorship ObfuscationAsad Mahmood, Zubair Shafiq, Padmini SrinivasanACL 2020 · 1 citation
- Revision in Continuous Space: Unsupervised Text Style Transfer without Adversarial LearningDayiheng Liu, Jie Fu, Yidan Zhang, Chris Pal et al.AAAI 2020 · 53 citations
