Self-Supervised Quality Estimation for Machine Translation
Yuanhang Zheng, Zhixing Tan, Meng Zhang, Mieradilijiang Maimaiti, Huanbo Luan, Maosong Sun, Qun Liu, Yang Liu
Abstract
Quality estimation (QE) of machine translation (MT) aims to evaluate the quality of machine-translated sentences without references and is important in practical applications of MT. Training QE models require massive parallel data with hand-crafted quality annotations, which are time-consuming and laborintensive to obtain. To address the issue of the absence of annotated training data, previous studies attempt to develop unsupervised QE methods. However, very few of them can be applied to both sentence-and word-level QE tasks, and they may suffer from noises in the synthetic data. To reduce the negative impact of noises, we propose a self-supervised method for both sentence-and word-level QE, which performs quality estimation by recovering the masked target words. Experimental results show that our method outperforms previous unsupervised methods on several QE tasks in different language pairs and domains. 1 * Corresponding author 1 Code can be found at https://github.com/ THUNLP-MT/SelfSupervisedQE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7d60a03b-cfae-46ee-9039-075429e9eb40Cited by top-tier papers3
- Improved Pseudo Data for Machine Translation Quality Estimation with Constrained Beam SearchXiang Geng, Yu Zhang, Zhejian Lai, Shuaijie She et al.EMNLP 2023 · 2 citations
- Case-Based Decision-Theoretic Decoding with Quality MemoriesHiroyuki Deguchi, Masaaki NagataEMNLP 2025
- Alleviating Distribution Shift in Synthetic Data for Machine Translation Quality EstimationXiang Geng, Zhejian Lai, Jiajun Chen, Hao Yang et al.ACL 2025
Builds on5
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- DirectQE: Direct Pretraining for Machine Translation Quality EstimationQu Cui, Shujian Huang, Jiahuan Li, Xiang Geng et al.AAAI 2021 · 24 citations
- Mask-Align: Self-Supervised Neural Word AlignmentChi Chen, Maosong Sun, Yang LiuACL 2021
- A Bidirectional Transformer Based Alignment Model for Unsupervised Word AlignmentJingyi Zhang, Josef van GenabithACL 2021
Related papers
- Improving Translation Quality Estimation with Bias MitigationHui Huang, Shuangzhi Wu, Kehai Chen, Hui Di et al.ACL 2023 · 2 citations
- Denoising Pre-training for Machine Translation Quality Estimation with Curriculum LearningXiang Geng, Yu Zhang, Jiahuan Li, Shujian Huang et al.AAAI 2023 · 11 citations
- Competency-Aware Neural Machine Translation: Can Machine Translation Know its Own Translation Quality?Pei Zhang, Baosong Yang, Haoran Wei, Dayiheng Liu et al.EMNLP 2022 · 1 citation
- Classification-based Quality Estimation: Small and Efficient Models for Real-world ApplicationsShuo Sun, Ahmed El-Kishky, Vishrav Chaudhary, James Cross et al.EMNLP 2021 · 1 citation
- Investigating the Helpfulness of Word-Level Quality Estimation for Post-Editing Machine Translation OutputRaksha Shenoy, Nico Herbig, Antonio Krüger, Josef van GenabithEMNLP 2021 · 3 citations
