Directed Acyclic Transformer for Non-Autoregressive Machine Translation
Fei Huang, Hao Zhou, Yang Liu, Hang Li, Minlie Huang
Abstract
Non-autoregressive Transformers (NATs) significantly reduce the decoding latency by generating all tokens in parallel. However, such independent predictions prevent NATs from capturing the dependencies between the tokens for generating multiple possible translations. In this paper, we propose Directed Acyclic Transfomer (DA-Transformer), which represents the hidden states in a Directed Acyclic Graph (DAG), where each path of the DAG corresponds to a specific translation. The whole DAG simultaneously captures multiple translations and facilitates fast predictions in a non-autoregressive fashion. Experiments on the raw training data of WMT benchmark show that DA-Transformer substantially outperforms previous NATs by about 3 BLEU on average, which is the first NAT model that achieves competitive results with autoregressive Transformers without relying on knowledge distillation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- On the Learning of Non-Autoregressive TransformersFei Huang, Tianhua Tao, Hao Zhou, Lei Li et al.ICML 2022 · 35 citations
- DASpeech: Directed Acyclic Transformer for Fast and High-quality Speech-to-Speech TranslationQingkai Fang, Yan Zhou, Yang FengNeurIPS 2023 · 22 citations
- Non-autoregressive Machine Translation with Probabilistic Context-free GrammarShangtong Gui, Chenze Shao, Zhengrui Ma, Xishan Zhang et al.NeurIPS 2023 · 16 citations
- AMOM: Adaptive Masking over Masking for Conditional Masked Language ModelYisheng Xiao, Ruiyang Xu, Lijun Wu, Juntao Li et al.AAAI 2023 · 14 citations
- Non-Autoregressive Math Word Problem Solver with Unified Tree StructureYi Bin, Mengqun Han, Wenhao Shi, Lei Wang et al.EMNLP 2023 · 7 citations
Builds on19
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Understanding Knowledge Distillation in Non-autoregressive Machine TranslationChunting Zhou, Jiatao Gu, Graham NeubigICLR 2020 · 235 citations
- Deep Encoder, Shallow Decoder: Reevaluating Non-autoregressive Machine TranslationJungo Kasai, Nikolaos Pappas, Hao Peng, James Cross et al.ICLR 2021 · 154 citations
- Latent-Variable Non-Autoregressive Neural Machine Translation with Deterministic Inference Using a Delta PosteriorRaphael Shu, Jason Lee, Hideki Nakayama, Kyunghyun ChoAAAI 2020 · 125 citations
- Aligned Cross Entropy for Non-Autoregressive Machine TranslationMarjan Ghazvininejad, Vladimir Karpukhin, Luke Zettlemoyer, Omer LevyICML 2020 · 121 citations
Related papers
- Fuzzy Alignments in Directed Acyclic Graph for Non-Autoregressive Machine TranslationZhengrui Ma, Chenze Shao, Shangtong Gui, Min Zhang et al.ICLR 2023 · 5 citations
- Non-autoregressive Machine Translation with Disentangled Context TransformerJungo Kasai, James Cross, Marjan Ghazvininejad, Jiatao GuICML 2020 · 113 citations
- Non-autoregressive Streaming Transformer for Simultaneous TranslationZhengrui Ma, Shaolei Zhang, Shoutao Guo, Chenze Shao et al.EMNLP 2023 · 3 citations
- Selective Knowledge Distillation for Non-Autoregressive Neural Machine TranslationMin Liu, Yu Bao, Chengqi Zhao, Shujian HuangAAAI 2023 · 4 citations
- NAT4AT: Using Non-Autoregressive Translation Makes Autoregressive Translation Faster and BetterHuanran Zheng, Wei Zhu, Xiaoling WangWWW 2024 · 13 citations
