Robustness to Programmable String Transformations via Augmented Abstract Training
Yuhao Zhang, Aws Albarghouthi, Loris D'Antoni
摘要
Deep neural networks for natural language processing tasks are vulnerable to adversarial input perturbations. In this paper, we present a versatile language for programmatically specifying string transformations -- e.g., insertions, deletions, substitutions, swaps, etc. -- that are relevant to the task at hand. We then present an approach to adversarially training models that are robust to such user-defined string transformations. Our approach combines the advantages of search-based techniques for adversarial training with abstraction-based techniques. Specifically, we show how to decompose a set of user-defined string transformations into two component specifications, one that benefits from search and another from abstraction. We use our technique to train models on the AG and SST2 datasets and show that the resulting models are robust to combinations of user-defined transformations mimicking spelling mistakes and other meaning-preserving transformations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Adversarial Robustness for CodePavol Bielik, Martin T. VechevICML 2020 · 被引用 101 次
- RoPGen: Towards Robust Code Authorship Attribution via Automatic Coding Style TransformationZhen Li, Qian (Guenevere) Chen, Chen Chen, Yayi Zou 等ICSE 2022 · 被引用 39 次
- Certifying Robustness to Programmable Data Bias in Decision TreesAnna P. Meyer, Aws Albarghouthi, Loris D'AntoniNeurIPS 2021 · 被引用 34 次
- Interval universal approximation for neural networksZi Wang, Aws Albarghouthi, Gautam Prakriya, Somesh JhaPOPL 2022 · 被引用 18 次
- Fast and precise certification of transformersGregory Bonaert, Dimitar I. Dimitrov, Maximilian Baader, Martin T. VechevPLDI 2021 · 被引用 18 次
它引用的顶会 Paper4
- Scalable Verified Training for Provably Robust Image ClassificationSven Gowal, Krishnamurthy Dvijotham, Robert Stanforth, Rudy Bunel 等ICCV 2019 · 被引用 196 次
- Adversarial Training and Provable Defenses: Bridging the GapMislav Balunovic, Martin T. VechevICLR 2020 · 被引用 186 次
- Robustness Verification for TransformersZhouxing Shi, Huan Zhang, Kai-Wei Chang, Minlie Huang 等ICLR 2020 · 被引用 131 次
- Towards Verified Robustness under Text Deletion InterventionsJohannes Welbl, Po-Sen Huang, Robert Stanforth, Sven Gowal 等ICLR 2020 · 被引用 7 次
相关 Paper
- Certified Robustness to Programmable Transformations in LSTMsYuhao Zhang, Aws Albarghouthi, Loris D'AntoniEMNLP 2021 · 被引用 8 次
- Word Level Robustness Enhancement: Fight Perturbation with PerturbationPei Huang, Yuting Yang, Fuqi Jia, Minghao Liu 等AAAI 2022 · 被引用 14 次
- Discrete Adversarial Attack to Models of CodeFengjuan Gao, Yu Wang, Ke WangPLDI 2023 · 被引用 23 次
- Attribute-Guided Adversarial Training for Robustness to Natural PerturbationsTejas Gokhale, Rushil Anirudh, Bhavya Kailkhura, Jayaraman J. Thiagarajan 等AAAI 2021 · 被引用 42 次
- RMLM: A Flexible Defense Framework for Proactively Mitigating Word-level Adversarial AttacksZhaoyang Wang, Zhiyue Liu, Xiaopeng Zheng, Qinliang Su 等ACL 2023 · 被引用 16 次
