Testing Machine Translation via Referential Transparency
Pinjia He, Clara Meister, Zhendong Su
摘要
Machine translation software has seen rapid progress in recent years due to the advancement of deep Neural Networks. People routinely use machine translation software in their daily lives for tasks such as ordering food in a foreign restaurant, receiving medical diagnosis and treatment from foreign doctors, and reading international political news online. However, due to the complexity and intractability of the underlying Neural Networks, modern machine translation software is still far from robust and can produce poor or incorrect translations; this can lead to misunderstanding, financial loss, threats to personal safety and health, and political conflicts. To address this problem, we introduce referentially transparent inputs (RTIs), a simple, widely applicable methodology for validating machine translation software. A referentially transparent input is a piece of text that should have similar translations when used in different contexts. Our practical implementation, Purity, detects when this property is broken by a translation. To evaluate RTI, we use Purity to test Google Translate and Bing Microsoft Translator with 200 unlabeled sentences, which detected 123 and 142 erroneous translations with high precision (79.3% and 78.3%). The translation errors are diverse, including examples of under-translation, over-translation, word/phrase mistranslation, incorrect modification, and unclear logic.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- BiasAsker: Measuring the Bias in Conversational AI SystemYuxuan Wan, Wenxuan Wang, Pinjia He, Jiazhen Gu 等FSE 2023 · 被引用 50 次
- CCTEST: Testing and Repairing Code Completion SystemsZongjie Li, Chaozheng Wang, Zhibo Liu, Haoxuan Wang 等ICSE 2023 · 被引用 49 次
- Improving Machine Translation Systems via Isotopic ReplacementZeyu Sun, Jie M. Zhang, Yingfei Xiong, Mark Harman 等ICSE 2022 · 被引用 42 次
- Automated testing of image captioning systemsBoxi Yu, Zhiqing Zhong, Xinran Qin, Jiayi Yao 等ISSTA 2022 · 被引用 24 次
- MTTM: Metamorphic Testing for Textual Content Moderation SoftwareWenxuan Wang, Jen-tse Huang, Weibin Wu, Jianping Zhang 等ICSE 2023 · 被引用 23 次
它引用的顶会 Paper5
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li 等NDSS 2019 · 被引用 876 次
- Hidden Voice CommandsNicholas Carlini, Pratyush Mishra, Tavish Vaidya, Yuankai Zhang 等USENIX Security 2016 · 被引用 672 次
- Automatic testing and improvement of machine translationZeyu Sun, Jie M. Zhang, Mark Harman, Mike Papadakis 等ICSE 2020 · 被引用 111 次
- Structure-invariant testing for machine translationPinjia He, Clara Meister, Zhendong SuICSE 2020 · 被引用 84 次
- Machine translation testing via pathological invarianceShashij Gupta, Pinjia He, Clara Meister, Zhendong SuFSE 2020 · 被引用 42 次
相关 Paper
- Automated Testing for Machine Translation via Constituency InvariancePin Ji, Yang Feng, Jia Liu, Zhihong Zhao 等ASE 2021 · 被引用 14 次
- Back Deduction Based Testing for Word Sense Disambiguation Ability of Machine Translation SystemsJun Wang, Yanhui Li, Xiang Huang, Lin Chen 等ISSTA 2023 · 被引用 4 次
- Evaluating Terminology Translation in Machine Translation Systems via Metamorphic TestingYihui Xu, Yanhui Li, Jun Wang, Xiaofang ZhangASE 2024 · 被引用 2 次
- Extrinsic Evaluation of Machine Translation MetricsNikita Moghe, Tom Sherborne, Mark Steedman, Alexandra BirchACL 2023 · 被引用 12 次
- Did Translation Models Get More Robust Without Anyone Even Noticing?Ben Peters, André F. T. MartinsACL 2025 · 被引用 10 次
