On the Learning of Non-Autoregressive Transformers
Fei Huang, Tianhua Tao, Hao Zhou, Lei Li, Minlie Huang
Abstract
Non-autoregressive Transformer (NAT) is a family of text generation models, which aims to reduce the decoding latency by predicting the whole sentences in parallel. However, such latency reduction sacrifices the ability to capture left-to-right dependencies, thereby making NAT learning very challenging. In this paper, we present theoretical and empirical analyses to reveal the challenges of NAT learning and propose a unified perspective to understand existing successes. First, we show that simply training NAT by maximizing the likelihood can lead to an approximation of marginal distributions but drops all dependencies between tokens, where the dropped information can be measured by the dataset's conditional total correlation. Second, we formalize many previous objectives in a unified framework and show that their success can be concluded as maximizing the likelihood on a proxy distribution, leading to a reduced information loss. Empirical studies show that our perspective can explain the phenomena in NAT learning and guide the design of new training methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- DiffuSeq: Sequence to Sequence Text Generation with Diffusion ModelsShansan Gong, Mukai Li, Jiangtao Feng, Zhiyong Wu et al.ICLR 2023 · 94 citations
- Directed Acyclic Transformer for Non-Autoregressive Machine TranslationFei Huang, Hao Zhou, Yang Liu, Hang Li et al.ICML 2022 · 82 citations
- Theoretical Benefit and Limitation of Diffusion Language ModelGuhao Feng, Yihan Geng, Jian Guan, Wei Wu et al.NeurIPS 2025 · 52 citations
- ParallelBench: Understanding the Trade-offs of Parallel Decoding in Diffusion LLMsWonjun Kang, Kevin Galim, Seunghyuk Oh, Minjae Lee et al.ICLR 2026 · 49 citations
- Transformers over Directed Acyclic GraphsYuankai Luo, Veronika Thost, Lei ShiNeurIPS 2023 · 43 citations
Builds on19
- Understanding Knowledge Distillation in Non-autoregressive Machine TranslationChunting Zhou, Jiatao Gu, Graham NeubigICLR 2020 · 235 citations
- A Theory of Usable Information under Computational ConstraintsYilun Xu, Shengjia Zhao, Jiaming Song, Russell Stewart et al.ICLR 2020 · 211 citations
- Deep Encoder, Shallow Decoder: Reevaluating Non-autoregressive Machine TranslationJungo Kasai, Nikolaos Pappas, Hao Peng, James Cross et al.ICLR 2021 · 154 citations
- Latent-Variable Non-Autoregressive Neural Machine Translation with Deterministic Inference Using a Delta PosteriorRaphael Shu, Jason Lee, Hideki Nakayama, Kyunghyun ChoAAAI 2020 · 125 citations
- Aligned Cross Entropy for Non-Autoregressive Machine TranslationMarjan Ghazvininejad, Vladimir Karpukhin, Luke Zettlemoyer, Omer LevyICML 2020 · 121 citations
Related papers
- AEQA-NAT : Adaptive End-to-end Quantization Alignment Training Framework for Non-autoregressive Machine TranslationXiangyu Qu, Guojing Liu, Liang LiICML 2025
- Non-autoregressive Streaming Transformer for Simultaneous TranslationZhengrui Ma, Shaolei Zhang, Shoutao Guo, Chenze Shao et al.EMNLP 2023 · 3 citations
- An EM Approach to Non-autoregressive Conditional Sequence GenerationZhiqing Sun, Yiming YangICML 2020 · 43 citations
- NAT4AT: Using Non-Autoregressive Translation Makes Autoregressive Translation Faster and BetterHuanran Zheng, Wei Zhu, Xiaoling WangWWW 2024 · 13 citations
- Multi-Granularity Optimization for Non-Autoregressive TranslationYafu Li, Leyang Cui, Yongjing Yin, Yue ZhangEMNLP 2022 · 9 citations
