Robustness to Programmable String Transformations via Augmented Abstract Training
Yuhao Zhang, Aws Albarghouthi, Loris D'Antoni
Abstract
Deep neural networks for natural language processing tasks are vulnerable to adversarial input perturbations. In this paper, we present a versatile language for programmatically specifying string transformations -- e.g., insertions, deletions, substitutions, swaps, etc. -- that are relevant to the task at hand. We then present an approach to adversarially training models that are robust to such user-defined string transformations. Our approach combines the advantages of search-based techniques for adversarial training with abstraction-based techniques. Specifically, we show how to decompose a set of user-defined string transformations into two component specifications, one that benefits from search and another from abstraction. We use our technique to train models on the AG and SST2 datasets and show that the resulting models are robust to combinations of user-defined transformations mimicking spelling mistakes and other meaning-preserving transformations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ebcc30c8-99a1-49a2-8f50-5a91a3da7996Cited by top-tier papers7
- Adversarial Robustness for CodePavol Bielik, Martin T. VechevICML 2020 · 101 citations
- RoPGen: Towards Robust Code Authorship Attribution via Automatic Coding Style TransformationZhen Li, Qian (Guenevere) Chen, Chen Chen, Yayi Zou et al.ICSE 2022 · 39 citations
- Certifying Robustness to Programmable Data Bias in Decision TreesAnna P. Meyer, Aws Albarghouthi, Loris D'AntoniNeurIPS 2021 · 34 citations
- Interval universal approximation for neural networksZi Wang, Aws Albarghouthi, Gautam Prakriya, Somesh JhaPOPL 2022 · 18 citations
- Fast and precise certification of transformersGregory Bonaert, Dimitar I. Dimitrov, Maximilian Baader, Martin T. VechevPLDI 2021 · 18 citations
Builds on4
- Scalable Verified Training for Provably Robust Image ClassificationSven Gowal, Krishnamurthy Dvijotham, Robert Stanforth, Rudy Bunel et al.ICCV 2019 · 196 citations
- Adversarial Training and Provable Defenses: Bridging the GapMislav Balunovic, Martin T. VechevICLR 2020 · 186 citations
- Robustness Verification for TransformersZhouxing Shi, Huan Zhang, Kai-Wei Chang, Minlie Huang et al.ICLR 2020 · 131 citations
- Towards Verified Robustness under Text Deletion InterventionsJohannes Welbl, Po-Sen Huang, Robert Stanforth, Sven Gowal et al.ICLR 2020 · 7 citations
Related papers
- Certified Robustness to Programmable Transformations in LSTMsYuhao Zhang, Aws Albarghouthi, Loris D'AntoniEMNLP 2021 · 8 citations
- Word Level Robustness Enhancement: Fight Perturbation with PerturbationPei Huang, Yuting Yang, Fuqi Jia, Minghao Liu et al.AAAI 2022 · 14 citations
- Discrete Adversarial Attack to Models of CodeFengjuan Gao, Yu Wang, Ke WangPLDI 2023 · 23 citations
- Attribute-Guided Adversarial Training for Robustness to Natural PerturbationsTejas Gokhale, Rushil Anirudh, Bhavya Kailkhura, Jayaraman J. Thiagarajan et al.AAAI 2021 · 42 citations
- RMLM: A Flexible Defense Framework for Proactively Mitigating Word-level Adversarial AttacksZhaoyang Wang, Zhiyue Liu, Xiaopeng Zheng, Qinliang Su et al.ACL 2023 · 16 citations
