Don't Retrain, Just Rewrite: Countering Adversarial Perturbations by Rewriting Text
Ashim Gupta, Carter Wood Blum, Temma Choji, Yingjie Fei, Shalin Shah, Alakananda Vempala, Vivek Srikumar
摘要
Can language models transform inputs to protect text classifiers against adversarial attacks? In this work, we present ATINTER, a model that intercepts and learns to rewrite adversarial inputs to make them non-adversarial for a downstream text classifier. Our experiments on four datasets and five attack mechanisms reveal that ATINTER is effective at providing better adversarial robustness than existing defense approaches, without compromising task accuracy. For example, on sentiment classification using the SST-2 dataset, our method improves the adversarial accuracy over the best existing defense approach by more than 4% with a smaller decrease in task accuracy (0.5 % vs. 2.5%). Moreover, we show that ATINTER generalizes across multiple downstream tasks and classifiers without having to explicitly retrain it for those settings. For example, we find that when ATINTER is trained to remove adversarial perturbations for the sentiment classification task on the SST-2 dataset, it even transfers to a semantically different task of news classification (on AGNews) and improves the adversarial robustness by more than 10%. BERT-SST-2 RoBERTa-SST-2 Aids warning over bushmeat barter: Meat from African wild animals ... in the UK is spreading a virus similar to HIV, a leading scientist warns. Aids warning over bushmeat trade: Meat from African wild animals ... in the UK is spreading a virus similar to HIV, a leading scientist warns. Business Adversarial input for BERT-SST-2 (incorrectly predicted positive) Adversarial input for RoBERTa-SST-2 (incorrectly predicted positive) Adversarial input for BERT-AGNews
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li 等NDSS 2019 · 被引用 876 次
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue 等EMNLP 2020 · 被引用 529 次
- Seq2Sick: Evaluating the Robustness of Sequence-to-Sequence Models with Adversarial ExamplesMinhao Cheng, Jinfeng Yi, Pin-Yu Chen, Huan Zhang 等AAAI 2020 · 被引用 268 次
相关 Paper
- Searching for an Effective Defender: Benchmarking Defense against Adversarial Word SubstitutionZongyi Li, Jianhan Xu, Jiehang Zeng, Linyang Li 等EMNLP 2021 · 被引用 46 次
- SHIELD: Defending Textual Neural Networks against Multiple Black-Box Adversarial Attacks with Stochastic Multi-Expert PatcherThai Le, Noseong Park, Dongwon LeeACL 2022 · 被引用 27 次
- Disentangled Information Bottleneck for Adversarial Text DefenseYidan Xu, Xinghao Yang, Wei Liu, Bao-di Liu 等EMNLP 2025
- Improving the Robustness of Transformer-based Large Language Models with Dynamic AttentionLujia Shen, Yuwen Pu, Shouling Ji, Changjiang Li 等NDSS 2024
- RMLM: A Flexible Defense Framework for Proactively Mitigating Word-level Adversarial AttacksZhaoyang Wang, Zhiyue Liu, Xiaopeng Zheng, Qinliang Su 等ACL 2023 · 被引用 16 次
