Fuzzy Alignments in Directed Acyclic Graph for Non-Autoregressive Machine Translation
Zhengrui Ma, Chenze Shao, Shangtong Gui, Min Zhang, Yang Feng
Abstract
Non-autoregressive translation (NAT) reduces the decoding latency but suffers from performance degradation due to the multi-modality problem. Recently, the structure of directed acyclic graph has achieved great success in NAT, which tackles the multi-modality problem by introducing dependency between vertices. However, training it with negative log-likelihood loss implicitly requires a strict alignment between reference tokens and vertices, weakening its ability to handle multiple translation modalities. In this paper, we hold the view that all paths in the graph are fuzzily aligned with the reference sentence. We do not require the exact alignment but train the model to maximize a fuzzy alignment score between the graph and reference, which takes captured translations in all modalities into account. Extensive experiments on major WMT benchmarks show that our method substantially improves translation performance and increases prediction confidence, setting a new state of the art for NAT on the raw training data. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c745dffd-2f44-4aa8-afbe-df351c1b9676Cited by top-tier papers9
- DASpeech: Directed Acyclic Transformer for Fast and High-quality Speech-to-Speech TranslationQingkai Fang, Yan Zhou, Yang FengNeurIPS 2023 · 22 citations
- Non-autoregressive Machine Translation with Probabilistic Context-free GrammarShangtong Gui, Chenze Shao, Zhengrui Ma, Xishan Zhang et al.NeurIPS 2023 · 16 citations
- Beyond MLE: Convex Learning for Text GenerationChenze Shao, Zhengrui Ma, Min Zhang, Yang FengNeurIPS 2023 · 5 citations
- Self-Improvement of Non-autoregressive Model via Sequence-Level DistillationYusheng Liao, Shuyang Jiang, Yiqi Li, Yu Wang et al.EMNLP 2023 · 4 citations
- Non-autoregressive Streaming Transformer for Simultaneous TranslationZhengrui Ma, Shaolei Zhang, Shoutao Guo, Chenze Shao et al.EMNLP 2023 · 3 citations
Builds on16
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Understanding Knowledge Distillation in Non-autoregressive Machine TranslationChunting Zhou, Jiatao Gu, Graham NeubigICLR 2020 · 235 citations
- Deep Encoder, Shallow Decoder: Reevaluating Non-autoregressive Machine TranslationJungo Kasai, Nikolaos Pappas, Hao Peng, James Cross et al.ICLR 2021 · 154 citations
- Aligned Cross Entropy for Non-Autoregressive Machine TranslationMarjan Ghazvininejad, Vladimir Karpukhin, Luke Zettlemoyer, Omer LevyICML 2020 · 121 citations
- Non-autoregressive Machine Translation with Disentangled Context TransformerJungo Kasai, James Cross, Marjan Ghazvininejad, Jiatao GuICML 2020 · 113 citations
Related papers
- Directed Acyclic Transformer for Non-Autoregressive Machine TranslationFei Huang, Hao Zhou, Yang Liu, Hang Li et al.ICML 2022 · 82 citations
- Order-Agnostic Cross Entropy for Non-Autoregressive Machine TranslationCunxiao Du, Zhaopeng Tu, Jing JiangICML 2021 · 93 citations
- Non-Monotonic Latent Alignments for CTC-Based Non-Autoregressive Machine TranslationChenze Shao, Yang FengNeurIPS 2022 · 26 citations
- Rephrasing the Reference for Non-autoregressive Machine TranslationChenze Shao, Jinchao Zhang, Jie Zhou, Yang FengAAAI 2023 · 6 citations
- Multi-Granularity Optimization for Non-Autoregressive TranslationYafu Li, Leyang Cui, Yongjing Yin, Yue ZhangEMNLP 2022 · 9 citations
