Cascaded Text Generation with Markov Transformers
Yuntian Deng, Alexander M. Rush
摘要
The two dominant approaches to neural text generation are fully autoregressive models, using serial beam search decoding, and non-autoregressive models, using parallel decoding with no output dependencies. This work proposes an autoregressive model with sub-linear parallel time generation. Noting that conditional random fields with bounded context can be decoded in parallel, we propose an efficient cascaded decoding approach for generating high-quality output. To parameterize this cascade, we introduce a Markov transformer, a variant of the popular fully autoregressive model that allows us to simultaneously decode with specific autoregressive context cutoffs. This approach requires only a small modification from standard autoregressive training, while showing competitive accuracy/speed tradeoff compared to existing methods on five machine translation datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Efficient Conformal Prediction via Cascaded Inference with Expanded AdmissionAdam Fisch, Tal Schuster, Tommi S. Jaakkola, Regina BarzilayICLR 2021 · 被引用 53 次
- A Character-Level Length-Control Algorithm for Non-Autoregressive Sentence SummarizationPuyuan Liu, Xiang Zhang, Lili MouNeurIPS 2022 · 被引用 21 次
- Glancing Transformer for Non-Autoregressive Neural Machine TranslationLihua Qian, Hao Zhou, Yu Bao, Mingxuan Wang 等ACL 2021
- Conformal Generative Modeling with Improved Sample Efficiency through Sequential Greedy FilteringKlaus-Rudolf Kladny, Bernhard Schölkopf, Michael MuehlebachICLR 2025
它引用的顶会 Paper4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan 等ICLR 2020 · 被引用 683 次
- Understanding Knowledge Distillation in Non-autoregressive Machine TranslationChunting Zhou, Jiatao Gu, Graham NeubigICLR 2020 · 被引用 235 次
- Non-Autoregressive Machine Translation with Latent AlignmentsChitwan Saharia, William Chan, Saurabh Saxena, Mohammad NorouziEMNLP 2020 · 被引用 6 次
相关 Paper
- Non-autoregressive Machine Translation with Disentangled Context TransformerJungo Kasai, James Cross, Marjan Ghazvininejad, Jiatao GuICML 2020 · 被引用 113 次
- Deep Encoder, Shallow Decoder: Reevaluating Non-autoregressive Machine TranslationJungo Kasai, Nikolaos Pappas, Hao Peng, James Cross 等ICLR 2021 · 被引用 154 次
- Accelerating Transformer Inference for Translation via Parallel DecodingAndrea Santilli, Silvio Severino, Emilian Postolache, Valentino Maiorca 等ACL 2023 · 被引用 19 次
- Directed Acyclic Transformer for Non-Autoregressive Machine TranslationFei Huang, Hao Zhou, Yang Liu, Hang Li 等ICML 2022 · 被引用 82 次
- An EM Approach to Non-autoregressive Conditional Sequence GenerationZhiqing Sun, Yiming YangICML 2020 · 被引用 43 次
