A Reinforced Generation of Adversarial Examples for Neural Machine Translation
Wei Zou, Shujian Huang, Jun Xie, Xinyu Dai, Jiajun Chen
Abstract
Neural machine translation systems tend to fail on less decent inputs despite its significant efficacy, which may significantly harm the credibility of these systems-fathoming how and when neural-based systems fail in such cases is critical for industrial maintenance. Instead of collecting and analyzing bad cases using limited handcrafted error features, here we investigate this issue by generating adversarial examples via a new paradigm based on reinforcement learning. Our paradigm could expose pitfalls for a given performance metric, e.g., BLEU, and could target any given neural machine translation architecture. We conduct experiments of adversarial attacks on two mainstream neural machine translation architectures, RNN-search, and Transformer. The results show that our method efficiently produces stable attacks with meaning-preserving adversarial examples. We also present a qualitative and quantitative analysis for the preference pattern of the attack, demonstrating its capability of pitfall exposure.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c5c1bb8c-5032-4bf5-a645-ab5fe1093aa2Cited by top-tier papers13
- LAS-AT: Adversarial Training with Learnable Attack StrategyXiaojun Jia, Yong Zhang, Baoyuan Wu, Ke Ma et al.CVPR 2022 · 140 citations
- Bad Characters: Imperceptible NLP AttacksNicholas Boucher, Ilia Shumailov, Ross Anderson, Nicolas PapernotS&P 2022 · 133 citations
- SSPAttack: A Simple and Sweet Paradigm for Black-Box Hard-Label Textual Adversarial AttackHan Liu, Zhi Xu, Xiaotong Zhang, Xiaoming Xu et al.AAAI 2023 · 31 citations
- NMTSloth: understanding and testing efficiency degradation of neural machine translation systemsSimin Chen, Cong Liu, Mirazul Haque, Zihe Song et al.FSE 2022 · 22 citations
- Character-level White-Box Adversarial Attacks against Transformers via Attachable Subwords SubstitutionAiwei Liu, Honghai Yu, Xuming Hu, Shu'ang Li et al.EMNLP 2022 · 20 citations
Builds on1
Related papers
- Crafting Adversarial Examples for Neural Machine TranslationXinze Zhang, Junzhe Zhang, Zhenhua Chen, Kun HeACL 2021
- Imitation Attacks and Defenses for Black-box Machine Translation SystemsEric Wallace, Mitchell Stern, Dawn SongEMNLP 2020 · 63 citations
- Breeding Machine Translations: Evolutionary approach to survive and thrive in the world of automated evaluationJosef Jon, Ondrej BojarACL 2023 · 3 citations
- AdvAug: Robust Adversarial Augmentation for Neural Machine TranslationYong Cheng, Lu Jiang, Wolfgang Macherey, Jacob EisensteinACL 2020 · 105 citations
- BLEURT Has Universal Translations: An Analysis of Automatic Metrics by Minimum Risk TrainingYiming Yan, Tao Wang, Chengqi Zhao, Shujian Huang et al.ACL 2023 · 7 citations
