Online Learning Meets Machine Translation Evaluation: Finding the Best Systems with the Least Human Effort
Vânia Mendonça, Ricardo Rei, Luísa Coheur, Alberto Sardinha, Ana Lúcia Santos
Abstract
In Machine Translation, assessing the quality of a large amount of automatic translations can be challenging. Automatic metrics are not reliable when it comes to high performing systems. In addition, resorting to human evaluators can be expensive, especially when evaluating multiple systems. To overcome the latter challenge, we propose a novel application of online learning that, given an ensemble of Machine Translation systems, dynamically converges to the best systems, by taking advantage of the human feedback available. Our experiments on WMT'19 datasets show that our online approach quickly converges to the top-3 ranked systems for the language pairs considered, despite the lack of human feedback for many translations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 947cc1ed-86df-4c2c-8ac4-aac7e96e279cCited by top-tier papers3
- Better than Random: Reliable NLG Human Evaluation with Constrained Active SamplingJie Ruan, Xiao Pu, Mingqi Gao, Xiaojun Wan et al.AAAI 2024 · 8 citations
- Multi-Source Test-Time Adaptation as Dueling Bandits for Extractive Question AnsweringHai Ye, Qizhe Xie, Hwee Tou NgACL 2023 · 2 citations
- Simulating Bandit Learning from User Feedback for Extractive Question AnsweringGe Gao, Eunsol Choi, Yoav ArtziACL 2022
Builds on2
Related papers
- Non-parametric Online Learning from Human Feedback for Neural Machine TranslationDongqi Wang, Haoran Wei, Zhirui Zhang, Shujian Huang et al.AAAI 2022 · 15 citations
- Transductive Ensemble Learning for Neural Machine TranslationYiren Wang, Lijun Wu, Yingce Xia, Tao Qin et al.AAAI 2020 · 29 citations
- Automatic Machine Translation Evaluation in Many Languages via Zero-Shot ParaphrasingBrian Thompson, Matt PostEMNLP 2020 · 7 citations
- BLEU might be Guilty but References are not InnocentMarkus Freitag, David Grangier, Isaac CaswellEMNLP 2020 · 13 citations
- MT-Ranker: Reference-free machine translation evaluation by inter-system rankingIbraheem Muhammad Moosa, Rui Zhang, Wenpeng YinICLR 2024 · 13 citations
