Decoding As Dynamic Programming For Recurrent Autoregressive Models
Najam Zaidi, Trevor Cohn, Gholamreza Haffari
摘要
Decoding in autoregressive models (ARMs) consists of searching for a high scoring output sequence under the trained model. Standard decoding methods, based on unidirectional greedy algorithm or beam search, are suboptimal due to error propagation and myopic decisions which do not account for future steps in the generation process. In this paper we present a novel decoding approach based on the method of auxiliary coordinates (Carreira-Perpinan & Wang, 2014) to address the aforementioned shortcomings. Our method introduces discrete variables for output tokens, and auxiliary continuous variables representing the states of the underlying ARM. The auxiliary variables lead to a factor graph approximation of the ARM, whose maximum a posteriori (MAP) solution is found exactly using dynamic programming. The MAP solution is then used to recreate an improved factor graph approximation of the ARM via updated auxiliary variables. We then extend our approach to decode in an ensemble of ARMs, possibly with different generation orders, which is out of reach for the standard unidirectional decoding algorithms. Experiments on the text infilling task over SWAG and Daily Dialogue datasets show that our decoding method is superior to strong competing decoding methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Blank Language ModelsTianxiao Shen, Victor Quach, Regina Barzilay, Tommi S. JaakkolaEMNLP 2020 · 被引用 8 次
- Twist Decoding: Diverse Generators Guide Each OtherJungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Hao Peng 等EMNLP 2022
相关 Paper
- Self-Speculative Decoding Accelerates Lossless Inference in Any-Order and Any-Subset Autoregressive ModelsGabe Guo, Stefano ErmonICLR 2026
- Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked DiffusionsJaeyeon Kim, Kulin Shah, Vasilis Kontonis, Sham M. Kakade 等ICML 2025
- LUGS: Latent-aware Guidance for Efficient Unmasking in Diffusion Large Language ModelsNuanqiao Shan, Kairong Han, Xinpeng Dong, Kun KuangICML 2026
- Enabling Arbitrary Translation Objectives with Adaptive Tree SearchWang Ling, Wojciech Stokowiec, Domenic Donato, Chris Dyer 等ICLR 2022
- If beam search is the answer, what was the question?Clara Meister, Ryan Cotterell, Tim VieiraEMNLP 2020 · 被引用 26 次
