Breeding Machine Translations: Evolutionary approach to survive and thrive in the world of automated evaluation
Josef Jon, Ondrej Bojar
Abstract
We propose a genetic algorithm (GA) based method for modifying n-best lists produced by a machine translation (MT) system. Our method offers an innovative approach to improving MT quality and identifying weaknesses in evaluation metrics. Using common GA operations (mutation and crossover) on a list of hypotheses in combination with a fitness function (an arbitrary MT metric), we obtain novel and diverse outputs with high metric scores. With a combination of multiple MT metrics as the fitness function, the proposed method leads to an increase in translation quality as measured by other held-out automatic metrics. With a single metric (including popular ones such as COMET) as the fitness function, we find blind spots and flaws in the metric. This allows for an automated search for adversarial examples in an arbitrary metric, without prior assumptions on the form of such example. As a demonstration of the method, we create datasets of adversarial examples and use them to show that reference-free COMET is substantially less robust than the reference-based version.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c8a065ba-1d0b-49b9-9765-c22fbdb1287aCited by top-tier papers1
Ask how each one uses itBuilds on11
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 1,143 citations
- Statistical Power and Translationese in Machine Translation EvaluationYvette Graham, Barry Haddow, Philipp KoehnEMNLP 2020 · 82 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
- If beam search is the answer, what was the question?Clara Meister, Ryan Cotterell, Tim VieiraEMNLP 2020 · 26 citations
- Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation MetricsNitika Mathur, Timothy Baldwin, Trevor CohnACL 2020 · 14 citations
Related papers
- Crafting Adversarial Examples for Neural Machine TranslationXinze Zhang, Junzhe Zhang, Zhenhua Chen, Kun HeACL 2021
- DEMETR: Diagnosing Evaluation Metrics for TranslationMarzena Karpinska, Nishant Raj, Katherine Thai, Yixiao Song et al.EMNLP 2022 · 18 citations
- QUEST: Quality-Aware Metropolis-Hastings Sampling for Machine TranslationGonçalo Rui Alves Faria, Sweta Agrawal, António Farinhas, Ricardo Rei et al.NeurIPS 2024 · 23 citations
- A Reinforced Generation of Adversarial Examples for Neural Machine TranslationWei Zou, Shujian Huang, Jun Xie, Xinyu Dai et al.ACL 2020 · 66 citations
- Extrinsic Evaluation of Machine Translation MetricsNikita Moghe, Tom Sherborne, Mark Steedman, Alexandra BirchACL 2023 · 12 citations
