Digging Errors in NMT: Evaluating and Understanding Model Errors from Partial Hypothesis Space
Jianhao Yan, Chenming Wu, Fandong Meng, Jie Zhou
摘要
Solid evaluation of neural machine translation (NMT) is key to its understanding and improvement. Current evaluation of an NMT system is usually built upon a heuristic decoding algorithm (e.g., beam search) and an evaluation metric assessing similarity between the translation and golden reference. However, this system-level evaluation framework is limited by evaluating only one best hypothesis and search errors brought by heuristic decoding algorithms. To better understand NMT models, we propose a novel evaluation protocol, which defines model errors with model’s ranking capability over hypothesis space. To tackle the problem of exponentially large space, we propose two approximation methods, top region evaluation along with an exact top-k decoding algorithm, which finds top-ranked hypotheses in the whole hypothesis space, and Monte Carlo sampling evaluation, which simulates hypothesis space from a broader perspective. To quantify errors, we define our NMT model errors by measuring distance between the hypothesis array ranked by the model and the ideally ranked hypothesis array. After confirming the strong correlation with human judgment, we apply our evaluation to various NMT benchmarks and model architectures. We show that the state-of-the-art Transformer models face serious ranking issues and only perform at the random chance level in the top region. We further analyze model errors on architectures with different depths and widths, as well as different data-augmentation techniques, showing how these factors affect model errors. Finally, we connect model errors with the search algorithms and provide interesting findings of beam search inductive bias and correlation with Minimum Bayes Risk (MBR) decoding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Shallow-to-Deep Training for Neural Machine TranslationBei Li, Ziyang Wang, Hui Liu, Yufan Jiang 等EMNLP 2020 · 被引用 41 次
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 被引用 40 次
- If beam search is the answer, what was the question?Clara Meister, Ryan Cotterell, Tim VieiraEMNLP 2020 · 被引用 26 次
- Multi-Unit Transformers for Neural Machine TranslationJianhao Yan, Fandong Meng, Jie ZhouEMNLP 2020 · 被引用 21 次
相关 Paper
- Understanding the Properties of Minimum Bayes Risk Decoding in Neural Machine TranslationMathias Müller, Rico SennrichACL 2021
- QUEST: Quality-Aware Metropolis-Hastings Sampling for Machine TranslationGonçalo Rui Alves Faria, Sweta Agrawal, António Farinhas, Ricardo Rei 等NeurIPS 2024 · 被引用 23 次
- Machine Translation Decoding beyond Beam SearchRémi Leblond, Jean-Baptiste Alayrac, Laurent Sifre, Miruna Pislar 等EMNLP 2021 · 被引用 6 次
- Quality-Aware Translation Models: Efficient Generation and Quality Estimation in a Single ModelChristian Tomani, David Vilar, Markus Freitag, Colin Cherry 等ACL 2024
- Uncertainty Determines the Adequacy of the Mode and the Tractability of Decoding in Sequence-to-Sequence ModelsFelix Stahlberg, Ilia Kulikov, Shankar KumarACL 2022 · 被引用 13 次
