Structure-invariant testing for machine translation
Pinjia He, Clara Meister, Zhendong Su
摘要
In recent years, machine translation software has increasingly been integrated into our daily lives. People routinely use machine translation for various applications, such as describing symptoms to a foreign doctor and reading political news in a foreign language. However, the complexity and intractability of neural machine translation (NMT) models that power modern machine translation make the robustness of these systems difficult to even assess, much less guarantee. Machine translation systems can return inferior results that lead to misunderstanding, medical misdiagnoses, threats to personal safety, or political conflicts. Despite its apparent importance, validating the robustness of machine translation systems is very difficult and has, therefore, been much under-explored. To tackle this challenge, we introduce structure-invariant testing (SIT), a novel metamorphic testing approach for validating machine translation software. Our key insight is that the translation results of "similar" source sentences should typically exhibit similar sentence structures. Specifically, SIT (1) generates similar source sentences by substituting one word in a given sentence with semantically similar, syntactically equivalent words; (2) represents sentence structure by syntax parse trees (obtained via constituency or dependency parsing); (3) reports sentence pairs whose structures differ quantitatively by more than some threshold. To evaluate SIT, we use it to test Google Translate and Bing Microsoft Translator with 200 source sentences as input, which led to 64 and 70 buggy issues with 69.5% and 70% top-1 accuracy, respectively. The translation errors are diverse, including under-translation, over-translation, incorrect modification, word/phrase mistranslation, and unclear logic.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper28
- Finding bugs in database systems via query partitioningManuel Rigger, Zhendong SuOOPSLA 2020 · 被引用 116 次
- Improving Human-AI Collaboration With Descriptions of AI BehaviorÁngel Alexander Cabrera, Adam Perer, Jason I. HongCSCW 2023 · 被引用 85 次
- Correlations between deep neural network model coverage criteria and model qualityShenao Yan, Guanhong Tao, Xuwei Liu, Juan Zhai 等FSE 2020 · 被引用 75 次
- AUTOTRAINER: An Automatic DNN Training Problem Detection and Repair SystemXiaoyu Zhang, Juan Zhai, Shiqing Ma, Chao ShenICSE 2021 · 被引用 62 次
- Testing Machine Translation via Referential TransparencyPinjia He, Clara Meister, Zhendong SuICSE 2021 · 被引用 50 次
它引用的顶会 Paper6
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha 等S&P 2016 · 被引用 3,275 次
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 被引用 1,633 次
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li 等NDSS 2019 · 被引用 876 次
- Hidden Voice CommandsNicholas Carlini, Pratyush Mishra, Tavish Vaidya, Yuankai Zhang 等USENIX Security 2016 · 被引用 672 次
相关 Paper
- Automated Testing for Machine Translation via Constituency InvariancePin Ji, Yang Feng, Jia Liu, Zhihong Zhao 等ASE 2021 · 被引用 14 次
- Machine translation testing via pathological invarianceShashij Gupta, Pinjia He, Clara Meister, Zhendong SuFSE 2020 · 被引用 42 次
- Automatic testing and improvement of machine translationZeyu Sun, Jie M. Zhang, Mark Harman, Mike Papadakis 等ICSE 2020 · 被引用 111 次
- Back Deduction Based Testing for Word Sense Disambiguation Ability of Machine Translation SystemsJun Wang, Yanhui Li, Xiang Huang, Lin Chen 等ISSTA 2023 · 被引用 4 次
- Improving Machine Translation Systems via Isotopic ReplacementZeyu Sun, Jie M. Zhang, Yingfei Xiong, Mark Harman 等ICSE 2022 · 被引用 42 次
