Defense against Synonym Substitution-based Adversarial Attacks via Dirichlet Neighborhood Ensemble
Yi Zhou, Xiaoqing Zheng, Cho-Jui Hsieh, Kai-Wei Chang, Xuanjing Huang
Abstract
Although deep neural networks have achieved prominent performance on many NLP tasks, they are vulnerable to adversarial examples. We propose Dirichlet Neighborhood Ensemble (DNE), a randomized method for training a robust model to defense synonym substitutionbased attacks. During training, DNE forms virtual sentences by sampling embedding vectors for each word in an input sentence from a convex hull spanned by the word and its synonyms, and it augments them with the training data. In such a way, the model is robust to adversarial attacks while maintaining the performance on the original clean data. DNE is agnostic to the network architectures and scales to large models (e.g., BERT) for NLP applications. Through extensive experimentation, we demonstrate that our method consistently outperforms recently proposed defense methods by a significant margin across different network architectures and multiple data sets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 92fe0425-0b71-400f-85e8-ad4fa1198ae8Cited by top-tier papers21
- TextHoaxer: Budgeted Hard-Label Adversarial Attacks on TextMuchao Ye, Chenglin Miao, Ting Wang, Fenglong MaAAAI 2022 · 57 citations
- Defending Against Alignment-Breaking Attacks via Robustly Aligned LLMBochuan Cao, Yuanpu Cao, Lu Lin, Jinghui ChenACL 2024 · 34 citations
- HQA-Attack: Toward High Quality Black-Box Hard-Label Adversarial Attack on TextHan Liu, Zhi Xu, Xiaotong Zhang, Feng Zhang et al.NeurIPS 2023 · 32 citations
- Adversarial Training for Improving Model Robustness? Look at Both Prediction and InterpretationHanjie Chen, Yangfeng JiAAAI 2022 · 31 citations
- SSPAttack: A Simple and Sweet Paradigm for Black-Box Hard-Label Textual Adversarial AttackHan Liu, Zhi Xu, Xiaotong Zhang, Xiaoming Xu et al.AAAI 2023 · 31 citations
Builds on6
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- Certified Robustness to Adversarial Examples with Differential PrivacyMathias Lécuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu et al.S&P 2019 · 1,022 citations
- Automatic Perturbation Analysis for Scalable Certified Robustness and BeyondKaidi Xu, Zhouxing Shi, Huan Zhang, Yihan Wang et al.NeurIPS 2020 · 415 citations
- Seq2Sick: Evaluating the Robustness of Sequence-to-Sequence Models with Adversarial ExamplesMinhao Cheng, Jinfeng Yi, Pin-Yu Chen, Huan Zhang et al.AAAI 2020 · 268 citations
- Robustness Verification for TransformersZhouxing Shi, Huan Zhang, Kai-Wei Chang, Minlie Huang et al.ICLR 2020 · 131 citations
Related papers
- Word Level Robustness Enhancement: Fight Perturbation with PerturbationPei Huang, Yuting Yang, Fuqi Jia, Minghao Liu et al.AAAI 2022 · 14 citations
- Searching for an Effective Defender: Benchmarking Defense against Adversarial Word SubstitutionZongyi Li, Jianhan Xu, Jiehang Zeng, Linyang Li et al.EMNLP 2021 · 46 citations
- Towards Robustness Against Natural Language Word SubstitutionsXinshuai Dong, Anh Tuan Luu, Rongrong Ji, Hong LiuICLR 2021 · 63 citations
- RMLM: A Flexible Defense Framework for Proactively Mitigating Word-level Adversarial AttacksZhaoyang Wang, Zhiyue Liu, Xiaopeng Zheng, Qinliang Su et al.ACL 2023 · 16 citations
- Certified Robustness to Programmable Transformations in LSTMsYuhao Zhang, Aws Albarghouthi, Loris D'AntoniEMNLP 2021 · 8 citations
