Training and Inference on Any-Order Autoregressive Models the Right Way
Andy Shih, Dorsa Sadigh, Stefano Ermon
摘要
Conditional inference on arbitrary subsets of variables is a core problem in probabilistic inference with important applications such as masked language modeling and image inpainting. In recent years, the family of Any-Order Autoregressive Models (AO-ARMs) -closely related to popular models such as BERT and XL-Net -has shown breakthrough performance in arbitrary conditional tasks across a sweeping range of domains. But, in spite of their success, in this paper we identify significant improvements to be made to previous formulations of AO-ARMs. First, we show that AO-ARMs suffer from redundancy in their probabilistic model, i.e., they define the same distribution in multiple different ways. We alleviate this redundancy by training on a smaller set of univariate conditionals that still maintains support for efficient arbitrary conditional inference. Second, we upweight the training loss for univariate conditionals that are evaluated more frequently during inference. Our method leads to improved performance with no compromises on tractability, giving state-of-the-art likelihoods in arbitrary conditional modeling on text (Text8), image (CIFAR10, ImageNet32), and continuous tabular data domains. 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang 等NeurIPS 2025 · 被引用 949 次
- Simplified and Generalized Masked Diffusion for Discrete DataJiaxin Shi, Kehang Han, Zhe Wang, Arnaud Doucet 等NeurIPS 2024 · 被引用 693 次
- Discrete Diffusion Modeling by Estimating the Ratios of the Data DistributionAaron Lou, Chenlin Meng, Stefano ErmonICML 2024 · 被引用 473 次
- Continuous Diffusion Model for Language ModelingJaehyeong Jo, Sung Ju HwangNeurIPS 2025 · 被引用 30 次
- Continuously Augmented Discrete Diffusion model for Categorical Generative ModelingHuangjie Zheng, Shansan Gong, Ruixiang Zhang, Tianrong Chen 等ICLR 2026 · 被引用 30 次
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow 等NeurIPS 2021 · 被引用 2,256 次
- Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingAlexander Ku, Peter Anderson, Roma Patel, Eugene Ie 等EMNLP 2020 · 被引用 208 次
相关 Paper
- Self-Speculative Decoding Accelerates Lossless Inference in Any-Order and Any-Subset Autoregressive ModelsGabe Guo, Stefano ErmonICLR 2026
- Autoregressive Models Rival Diffusion Models at ANY-ORDER GenerationTianqi Du, Lizhe Fang, Weijie Yang, Chenheng Zhang 等ICLR 2026 · 被引用 3 次
- Probabilistically Masked Language Model Capable of Autoregressive Generation in Arbitrary Word OrderYi Liao, Xin Jiang, Qun LiuACL 2020 · 被引用 28 次
- Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked DiffusionsJaeyeon Kim, Kulin Shah, Vasilis Kontonis, Sham M. Kakade 等ICML 2025
- Diverse Image Inpainting with Bidirectional and Autoregressive TransformersYingchen Yu, Fangneng Zhan, Rongliang Wu, Jianxiong Pan 等ACM MM 2021 · 被引用 153 次
