WeTS: A Benchmark for Translation Suggestion
Zhen Yang, Fandong Meng, Yingxue Zhang, Ernan Li, Jie Zhou
摘要
Translation suggestion (TS), which provides alternatives for specific words or phrases given the entire documents generated by machine translation (MT), has been proven to play a significant role in post-editing (PE). There are two main pitfalls for existing researches in this line. First, most conventional works only focus on the overall performance of PE but ignore the exact performance of TS, which makes the progress of PE sluggish and less explainable; Second, as no publicly available golden dataset exists to support in-depth research for TS, almost all of the previous works conduct experiments on their in-house datasets or the noisy datasets built automatically, which makes their experiments hard to be reproduced and compared. To break these limitations mentioned above and spur the research in TS, we create a benchmark dataset, called WeTS, which is a golden corpus annotated by expert translators on four translation directions. Apart from the golden corpus, we also propose several methods to generate synthetic corpora which can be used to improve the performance substantially through pre-training. As for the model, we propose the segment-aware self-attention based Transformer for TS. Experimental results show that our approach achieves the best results on all four directions, including Englishto-German, German-to-English, Chinese-to-English, and English-to-Chinese. Codes and corpus can be found at https://github.com/ ZhenYangIACAS/WeTS.git .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Cross-Align: Modeling Deep Cross-lingual Interactions for Word AlignmentSiyu Lai, Zhen Yang, Fandong Meng, Yufeng Chen 等EMNLP 2022 · 被引用 6 次
- Bilingual Synchronization: Restoring Translational Relationships with Editing OperationsJitao Xu, Josep Maria Crego, François YvonEMNLP 2022 · 被引用 5 次
它引用的顶会 Paper7
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 被引用 40 次
- Generating Diverse Translation by Manipulating Multi-Head AttentionZewei Sun, Shujian Huang, Hao-Ran Wei, Xinyu Dai 等AAAI 2020 · 被引用 36 次
- Neural Machine Translation Quality and Post-Editing PerformanceVilém Zouhar, Martin Popel, Ondrej Bojar, Ales TamchynaEMNLP 2021 · 被引用 13 次
- Generating Diverse Translation from Model Distribution with DropoutXuanfu Wu, Yang Feng, Chenze ShaoEMNLP 2020 · 被引用 10 次
相关 Paper
- Can Automatic Post-Editing Improve NMT?Shamil Chollampatt, Raymond Hendy Susanto, Liling Tan, Ewa SzymanskaEMNLP 2020 · 被引用 1 次
- Easy Guided Decoding in Providing Suggestions for Interactive Machine TranslationKe Wang, Xin Ge, Jiayi Wang, Yuqi Zhang 等ACL 2023
- GWLAN: General Word-Level AutocompletioN for Computer-Aided TranslationHuayang Li, Lemao Liu, Guoping Huang, Shuming ShiACL 2021
- LangMark: A Multilingual Dataset for Automatic Post-EditingDiego Velazquez, Mikaela Grace, Konstantinos Karageorgos, Lawrence Carin 等ACL 2025
- BTS: A Bi-lingual Benchmark for Text Segmentation in the WildXixi Xu, Zhongang Qi, Jianqi Ma, Honglun Zhang 等CVPR 2022 · 被引用 14 次
