Automatic Machine Translation Evaluation in Many Languages via Zero-Shot Paraphrasing
Brian Thompson, Matt Post
摘要
We frame the task of machine translation evaluation as one of scoring machine translation output with a sequence-to-sequence paraphraser, conditioned on a human reference. We propose training the paraphraser as a multilingual NMT system, treating paraphrasing as a zero-shot translation task (e.g., Czech to Czech). This results in the paraphraser's output mode being centered around a copy of the input sequence, which represents the best case scenario where the MT system output matches a human reference. Our method is simple and intuitive, and does not require human judgements for training. Our single model (trained in 39 languages) outperforms or statistically ties with all prior metrics on the WMT 2019 segment-level shared metrics task in all languages (excluding Gujarati where the model had no training data). We also explore using our model for the task of quality estimation as a metric-conditioning on the source instead of the reference-and find that it significantly outperforms every submission to the WMT 2019 shared task on quality estimation in every language pair. Word-level paraphraser log probabilities H(out|in) sBLEU LASER Copy Jason went to school at the University of Madrid . <EOS> -0.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper43
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 被引用 1,143 次
- COLD-Attack: Jailbreaking LLMs with Stealthiness and ControllabilityXingang Guo, Fangxu Yu, Huan Zhang, Lianhui Qin 等ICML 2024 · 被引用 173 次
- INSTRUCTSCORE: Towards Explainable Text Generation Evaluation with Automatic FeedbackWenda Xu, Danqing Wang, Liangming Pan, Zhenqiao Song 等EMNLP 2023 · 被引用 36 次
- ROSCOE: A Suite of Metrics for Scoring Step-by-Step ReasoningOlga Golovneva, Moya Chen, Spencer Poff, Martin Corredor 等ICLR 2023 · 被引用 28 次
- On the Limitations of Reference-Free Evaluations of Generated TextDaniel Deutsch, Rotem Dror, Dan RothEMNLP 2022 · 被引用 23 次
它引用的顶会 Paper3
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 被引用 40 次
- Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation MetricsNitika Mathur, Timothy Baldwin, Trevor CohnACL 2020 · 被引用 14 次
相关 Paper
- BLEU might be Guilty but References are not InnocentMarkus Freitag, David Grangier, Isaac CaswellEMNLP 2020 · 被引用 13 次
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 被引用 6 次
- Translationese as a Language in "Multilingual" NMTParker Riley, Isaac Caswell, Markus Freitag, David GrangierACL 2020
- Pre-training via ParaphrasingMike Lewis, Marjan Ghazvininejad, Gargi Ghosh, Armen Aghajanyan 等NeurIPS 2020 · 被引用 165 次
- MT-Ranker: Reference-free machine translation evaluation by inter-system rankingIbraheem Muhammad Moosa, Rui Zhang, Wenpeng YinICLR 2024 · 被引用 13 次
