WeTS: A Benchmark for Translation Suggestion
Zhen Yang, Fandong Meng, Yingxue Zhang, Ernan Li, Jie Zhou
Abstract
Translation suggestion (TS), which provides alternatives for specific words or phrases given the entire documents generated by machine translation (MT), has been proven to play a significant role in post-editing (PE). There are two main pitfalls for existing researches in this line. First, most conventional works only focus on the overall performance of PE but ignore the exact performance of TS, which makes the progress of PE sluggish and less explainable; Second, as no publicly available golden dataset exists to support in-depth research for TS, almost all of the previous works conduct experiments on their in-house datasets or the noisy datasets built automatically, which makes their experiments hard to be reproduced and compared. To break these limitations mentioned above and spur the research in TS, we create a benchmark dataset, called WeTS, which is a golden corpus annotated by expert translators on four translation directions. Apart from the golden corpus, we also propose several methods to generate synthetic corpora which can be used to improve the performance substantially through pre-training. As for the model, we propose the segment-aware self-attention based Transformer for TS. Experimental results show that our approach achieves the best results on all four directions, including Englishto-German, German-to-English, Chinese-to-English, and English-to-Chinese. Codes and corpus can be found at https://github.com/ ZhenYangIACAS/WeTS.git .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bf64739e-8d20-415b-ab1b-1f1c75d6f190Cited by top-tier papers2
- Cross-Align: Modeling Deep Cross-lingual Interactions for Word AlignmentSiyu Lai, Zhen Yang, Fandong Meng, Yufeng Chen et al.EMNLP 2022 · 6 citations
- Bilingual Synchronization: Restoring Translational Relationships with Editing OperationsJitao Xu, Josep Maria Crego, François YvonEMNLP 2022 · 5 citations
Builds on7
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
- Generating Diverse Translation by Manipulating Multi-Head AttentionZewei Sun, Shujian Huang, Hao-Ran Wei, Xinyu Dai et al.AAAI 2020 · 36 citations
- Neural Machine Translation Quality and Post-Editing PerformanceVilém Zouhar, Martin Popel, Ondrej Bojar, Ales TamchynaEMNLP 2021 · 13 citations
- Generating Diverse Translation from Model Distribution with DropoutXuanfu Wu, Yang Feng, Chenze ShaoEMNLP 2020 · 10 citations
Related papers
- Can Automatic Post-Editing Improve NMT?Shamil Chollampatt, Raymond Hendy Susanto, Liling Tan, Ewa SzymanskaEMNLP 2020 · 1 citation
- Easy Guided Decoding in Providing Suggestions for Interactive Machine TranslationKe Wang, Xin Ge, Jiayi Wang, Yuqi Zhang et al.ACL 2023
- GWLAN: General Word-Level AutocompletioN for Computer-Aided TranslationHuayang Li, Lemao Liu, Guoping Huang, Shuming ShiACL 2021
- LangMark: A Multilingual Dataset for Automatic Post-EditingDiego Velazquez, Mikaela Grace, Konstantinos Karageorgos, Lawrence Carin et al.ACL 2025
- BTS: A Bi-lingual Benchmark for Text Segmentation in the WildXixi Xu, Zhongang Qi, Jianqi Ma, Honglun Zhang et al.CVPR 2022 · 14 citations
