Improved Pseudo Data for Machine Translation Quality Estimation with Constrained Beam Search
Xiang Geng, Yu Zhang, Zhejian Lai, Shuaijie She, Wei Zou, Shimin Tao, Hao Yang, Jiajun Chen, Shujian Huang
Abstract
Machine translation (MT) quality estimation (QE) is a crucial task to estimate the quality of MT outputs when reference translations are unavailable. Many studies focus on generating pseudo data using large parallel corpus and achieve remarkable success in the supervised setting. However, pseudo data solutions are less satisfying in unsupervised scenarios because the pseudo labels are inaccurate or the pseudo translations differ from the real ones. To address these problems, we propose to generate pseudo data using the MT model with constrained beam search (CBSQE). CBSQE preserves the reference parts with high MT probabilities as correct translations, while the rest parts as the wrong ones for MT generation. Therefore, CBSQE can reduce the false negative labels caused by synonyms. Overall, beam search will prefer a more real hypothesis with a higher MT generation likelihood. Extensive experiments demonstrate that CBSQE outperforms strong baselines in both supervised and unsupervised settings. Analyses further show the superiority of CBSQE. The code is available at https://github.com/NJUNLP/njuqe .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e2bd5b7d-7641-4f9a-8344-356faf8d3624Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Understanding Knowledge Distillation in Non-autoregressive Machine TranslationChunting Zhou, Jiatao Gu, Graham NeubigICLR 2020 · 235 citations
- DirectQE: Direct Pretraining for Machine Translation Quality EstimationQu Cui, Shujian Huang, Jiahuan Li, Xiang Geng et al.AAAI 2021 · 24 citations
Related papers
- Self-Supervised Quality Estimation for Machine TranslationYuanhang Zheng, Zhixing Tan, Meng Zhang, Mieradilijiang Maimaiti et al.EMNLP 2021 · 5 citations
- Denoising Pre-training for Machine Translation Quality Estimation with Curriculum LearningXiang Geng, Yu Zhang, Jiahuan Li, Shujian Huang et al.AAAI 2023 · 11 citations
- Improving Translation Quality Estimation with Bias MitigationHui Huang, Shuangzhi Wu, Kehai Chen, Hui Di et al.ACL 2023 · 2 citations
- Meta Back-TranslationHieu Pham, Xinyi Wang, Yiming Yang, Graham NeubigICLR 2021 · 26 citations
- Competency-Aware Neural Machine Translation: Can Machine Translation Know its Own Translation Quality?Pei Zhang, Baosong Yang, Haoran Wei, Dayiheng Liu et al.EMNLP 2022 · 1 citation
