Adversarial Authorship Attribution for Deobfuscation
Wanyue Zhai, Jonathan Rusert, Zubair Shafiq, Padmini Srinivasan
摘要
Recent advances in natural language processing have enabled powerful privacy-invasive authorship attribution. To counter authorship attribution, researchers have proposed a variety of rule-based and learning-based text obfuscation approaches. However, existing authorship obfuscation approaches do not consider the adversarial threat model. Specifically, they are not evaluated against adversarially trained authorship attributors that are aware of potential obfuscation. To fill this gap, we investigate the problem of adversarial authorship attribution for deobfuscation. We show that adversarially trained authorship attributors are able to degrade the effectiveness of existing obfuscators from 20-30% to 5-10%. We also evaluate the effectiveness of adversarial training when the attributor makes incorrect assumptions about whether and which obfuscator was used. While there is a a clear degradation in attribution accuracy, it is noteworthy that this degradation is still at or above the attribution accuracy of the attributor that is not adversarially trained at all. Our results motivate the need to develop authorship obfuscation approaches that are resistant to deobfuscation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li 等NDSS 2019 · 被引用 876 次
- A4NT: Author Attribute Anonymity by Adversarial Training of Neural Machine TranslationRakshith Shetty, Bernt Schiele, Mario FritzUSENIX Security 2018 · 被引用 104 次
- Style Pooling: Automatic Text Style Obfuscation for Improved Classification FairnessFatemehsadat Mireshghallah, Taylor Berg-KirkpatrickEMNLP 2021 · 被引用 6 次
- A Girl Has A Name: Detecting Authorship ObfuscationAsad Mahmood, Zubair Shafiq, Padmini SrinivasanACL 2020 · 被引用 1 次
- Efficient Adversarial Training With Transferable Adversarial ExamplesHaizhong Zheng, Ziqi Zhang, Juncheng Gu, Honglak Lee 等CVPR 2020
相关 Paper
- ALISON: Fast and Effective Stylometric Authorship ObfuscationEric Xing, Saranya Venkatraman, Thai Le, Dongwon LeeAAAI 2024 · 被引用 7 次
- Misleading Authorship Attribution of Source Code using Adversarial LearningErwin Quiring, Alwin Maier, Konrad RieckUSENIX Security 2019 · 被引用 123 次
- A Multifaceted Framework to Evaluate Evasion, Content Preservation, and Misattribution in Authorship Obfuscation TechniquesMalik H. Altakrori, Thomas Scialom, Benjamin C. M. Fung, Jackie Chi Kit CheungEMNLP 2022 · 被引用 2 次
- "That Is a Suspicious Reaction!": Interpreting Logits Variation to Detect NLP Adversarial AttacksEdoardo Mosca, Shreyash Agarwal, Javier Rando-Ramirez, Georg GrohACL 2022 · 被引用 43 次
- Trade-offs and Guarantees of Adversarial Representation Learning for Information ObfuscationHan Zhao, Jianfeng Chi, Yuan Tian, Geoffrey J. GordonNeurIPS 2020 · 被引用 29 次
