PreQuEL: Quality Estimation of Machine Translation Outputs in Advance
Shachar Don-Yehiya, Leshem Choshen, Omri Abend
Abstract
We present the task of PreQuEL, Pre-(Quality-Estimation) Learning. A PreQuEL system predicts how well a given sentence will be translated, without recourse to the actual translation, thus eschewing unnecessary resource allocation when translation quality is bound to be low. PreQuEL can be defined relative to a given MT system (e.g., some industry service) or generally relative to the state-of-theart. From a theoretical perspective, PreQuEL places the focus on the source text, tracing properties, possibly linguistic features, that make a sentence harder to machine translate. We develop a baseline model for the task and analyze its performance. We also develop a data augmentation method (from parallel corpora), that improves results substantially. We show that this augmentation method can improve the performance of the Quality-Estimation task as well. 1 We investigate the properties of the input text that our model is sensitive to, by testing it on challenge sets and different languages. We conclude that it is aware of syntactic and semantic distinctions, and correlates and even over-emphasizes the importance of standard NLP features.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 353b0b3b-ca6f-48b5-81fa-db9f723b7847Cited by top-tier papers2
- Human Learning by Model Feedback: The Dynamics of Iterative Prompting with MidjourneyShachar Don-Yehiya, Leshem Choshen, Omri AbendEMNLP 2023 · 11 citations
- Where to start? Analyzing the potential value of intermediate modelsLeshem Choshen, Elad Venezian, Shachar Don-Yehiya, Noam Slonim et al.EMNLP 2023 · 7 citations
Builds on6
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Muppet: Massive Multi-task Representations with Pre-FinetuningArmen Aghajanyan, Anchit Gupta, Akshat Shrivastava, Xilun Chen et al.EMNLP 2021 · 176 citations
- Statistical Power and Translationese in Machine Translation EvaluationYvette Graham, Barry Haddow, Philipp KoehnEMNLP 2020 · 82 citations
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 6 citations
Related papers
- DirectQE: Direct Pretraining for Machine Translation Quality EstimationQu Cui, Shujian Huang, Jiahuan Li, Xiang Geng et al.AAAI 2021 · 24 citations
- Bias Mitigation in Machine Translation Quality EstimationHanna Behnke, Marina Fomicheva, Lucia SpeciaACL 2022
- SpeechQE: Estimating the Quality of Direct Speech TranslationHyoJung Han, Kevin Duh, Marine CarpuatEMNLP 2024 · 1 citation
- Revisiting Machine Translation for Cross-lingual ClassificationMikel Artetxe, Vedanuj Goswami, Shruti Bhosale, Angela Fan et al.EMNLP 2023 · 10 citations
- Denoising Pre-training for Machine Translation Quality Estimation with Curriculum LearningXiang Geng, Yu Zhang, Jiahuan Li, Shujian Huang et al.AAAI 2023 · 11 citations
