Discovering Non-monotonic Autoregressive Orderings with Variational Inference
Xuanlin Li, Brandon Trabucco, Dong Huk Park, Michael Luo, Sheng Shen, Trevor Darrell, Yang Gao
摘要
The predominant approach for language modeling is to process sequences from left to right, but this eliminates a source of information: the order by which the sequence was generated. One strategy to recover this information is to decode both the content and ordering of tokens. Existing approaches supervise content and ordering by designing problem-specific loss functions and pre-training with an ordering pre-selected. Other recent works use iterative search to discover problem-specific orderings for training, but suffer from high time complexity and cannot be efficiently parallelized. We address these limitations with an unsupervised parallelizable learner that discovers high-quality generation orders purely from training data -- no domain knowledge required. The learner contains an encoder network and decoder language model that perform variational inference with autoregressive orders (represented as permutation matrices) as latent variables. The corresponding ELBO is not differentiable, so we develop a practical algorithm for end-to-end optimization using policy gradients. We implement the encoder as a Transformer with non-causal attention that outputs permutations in one forward pass. Permutations then serve as target generation orders for training an insertion-based Transformer language model. Empirical results in language modeling tasks demonstrate that our method is context-aware and discovers orderings that are competitive with or even better than fixed orders.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Simple and Effective Masked Diffusion Language ModelsSubham S. Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan 等NeurIPS 2024 · 被引用 929 次
- The Pitfalls of Next-Token PredictionGregor Bachmann, Vaishnavh NagarajanICML 2024 · 被引用 163 次
- Generative Flow Networks for Discrete Probabilistic ModelingDinghuai Zhang, Nikolay Malkin, Zhen Liu, Alexandra Volokhova 等ICML 2022 · 被引用 131 次
- Training and Inference on Any-Order Autoregressive Models the Right WayAndy Shih, Dorsa Sadigh, Stefano ErmonNeurIPS 2022 · 被引用 68 次
- Improving Discrete Diffusion Unmasking Policies Beyond Explicit Reference PoliciesChunsan Hong, Seonho An, Min-Soo Kim, Jong Chul YeICLR 2026 · 被引用 23 次
它引用的顶会 Paper1
相关 Paper
- An Empirical Study of Generation Order for Machine TranslationWilliam Chan, Mitchell Stern, Jamie Kiros, Jakob UszkoreitEMNLP 2020 · 被引用 7 次
- Learning-Order Autoregressive Models with Application to Molecular Graph GenerationZhe Wang, Jiaxin Shi, Nicolas Heess, Arthur Gretton 等ICML 2025
- Unifying Masked Diffusion Models with Various Generation Orders and BeyondChunsan Hong, Sanghyun Lee, Jong Chul YEICML 2026
- Reinforced Context Order Recovery for Adaptive Reasoning and PlanningLong Ma, Fangwei Zhong, Yizhou WangNeurIPS 2025 · 被引用 4 次
- Variational Learning for Insertion-based GenerationYangtian Zhang, Zhe Wang, Arthur Gretton, ZHITAO YING 等ICML 2026
