Can Automatic Post-Editing Improve NMT?
Shamil Chollampatt, Raymond Hendy Susanto, Liling Tan, Ewa Szymanska
Abstract
Automatic post-editing (APE) aims to improve machine translations, thereby reducing human post-editing effort. APE has had notable success when used with statistical machine translation (SMT) systems but has not been as successful over neural machine translation (NMT) systems. This has raised questions on the relevance of APE task in the current scenario. However, the training of APE models has been heavily reliant on large-scale artificial corpora combined with only limited human post-edited data. We hypothesize that APE models have been underperforming in improving NMT translations due to the lack of adequate supervision. To ascertain our hypothesis, we compile a larger corpus of human post-edits of English to German NMT. We empirically show that a state-of-art neural APE model trained on this corpus can significantly improve a strong in-domain NMT system, challenging the current understanding in the field. We further investigate the effects of varying training data sizes, using artificial training data, and domain specificity for the APE task. We release this new corpus under CC BY-NC-SA 4.0 license at https:// github.com/shamilcm/pedra.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2bcb44b4-e2ec-4f9a-a2ce-73009f4a1e22Cited by top-tier papers2
- Ladder: A Model-Agnostic Framework Boosting LLM-based Machine Translation to the Next LevelZhaopeng Feng, Ruizhe Chen, Yan Zhang, Zijie Meng et al.EMNLP 2024 · 4 citations
- Reference Matters: Benchmarking Factual Error Correction for Dialogue Summarization with Fine-grained Evaluation FrameworkMingqi Gao, Xiaojun Wan, Jia Su, Zhefeng Wang et al.ACL 2023 · 4 citations
Related papers
- LangMark: A Multilingual Dataset for Automatic Post-EditingDiego Velazquez, Mikaela Grace, Konstantinos Karageorgos, Lawrence Carin et al.ACL 2025
- WeTS: A Benchmark for Translation SuggestionZhen Yang, Fandong Meng, Yingxue Zhang, Ernan Li et al.EMNLP 2022 · 5 citations
- Cross-lingual neural fuzzy matching for exploiting target-language monolingual corpora in computer-aided translationMiquel Esplà-Gomis, Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz, Felipe Sánchez-MartínezEMNLP 2022 · 2 citations
- Neural Machine Translation Quality and Post-Editing PerformanceVilém Zouhar, Martin Popel, Ondrej Bojar, Ales TamchynaEMNLP 2021 · 13 citations
- CURE: Code-Aware Neural Machine Translation for Automatic Program RepairNan Jiang, Thibaud Lutellier, Lin TanICSE 2021 · 267 citations
