Breeding Machine Translations: Evolutionary approach to survive and thrive in the world of automated evaluation
Josef Jon, Ondrej Bojar
摘要
We propose a genetic algorithm (GA) based method for modifying n-best lists produced by a machine translation (MT) system. Our method offers an innovative approach to improving MT quality and identifying weaknesses in evaluation metrics. Using common GA operations (mutation and crossover) on a list of hypotheses in combination with a fitness function (an arbitrary MT metric), we obtain novel and diverse outputs with high metric scores. With a combination of multiple MT metrics as the fitness function, the proposed method leads to an increase in translation quality as measured by other held-out automatic metrics. With a single metric (including popular ones such as COMET) as the fitness function, we find blind spots and flaws in the metric. This allows for an automated search for adversarial examples in an arbitrary metric, without prior assumptions on the form of such example. As a demonstration of the method, we create datasets of adversarial examples and use them to show that reference-free COMET is substantially less robust than the reference-based version.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 被引用 1,143 次
- Statistical Power and Translationese in Machine Translation EvaluationYvette Graham, Barry Haddow, Philipp KoehnEMNLP 2020 · 被引用 82 次
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 被引用 40 次
- If beam search is the answer, what was the question?Clara Meister, Ryan Cotterell, Tim VieiraEMNLP 2020 · 被引用 26 次
- Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation MetricsNitika Mathur, Timothy Baldwin, Trevor CohnACL 2020 · 被引用 14 次
相关 Paper
- Crafting Adversarial Examples for Neural Machine TranslationXinze Zhang, Junzhe Zhang, Zhenhua Chen, Kun HeACL 2021
- DEMETR: Diagnosing Evaluation Metrics for TranslationMarzena Karpinska, Nishant Raj, Katherine Thai, Yixiao Song 等EMNLP 2022 · 被引用 18 次
- QUEST: Quality-Aware Metropolis-Hastings Sampling for Machine TranslationGonçalo Rui Alves Faria, Sweta Agrawal, António Farinhas, Ricardo Rei 等NeurIPS 2024 · 被引用 23 次
- A Reinforced Generation of Adversarial Examples for Neural Machine TranslationWei Zou, Shujian Huang, Jun Xie, Xinyu Dai 等ACL 2020 · 被引用 66 次
- Extrinsic Evaluation of Machine Translation MetricsNikita Moghe, Tom Sherborne, Mark Steedman, Alexandra BirchACL 2023 · 被引用 12 次
