Learning-Order Autoregressive Models with Application to Molecular Graph Generation
Zhe Wang, Jiaxin Shi, Nicolas Heess, Arthur Gretton, Michalis K. Titsias
Abstract
Autoregressive models (ARMs) have become the workhorse for sequence generation tasks, since many problems can be modeled as next-token prediction. While there appears to be a natural ordering for text (i.e., left-to-right), for many data types, such as graphs, the canonical ordering is less obvious. To address this problem, we introduce a variant of ARM that generates high-dimensional data using a probabilistic ordering that is sequentially inferred from data. This model incorporates a trainable probability distribution, referred to as an order-policy, that dynamically decides the autoregressive order in a state-dependent manner. To train the model, we introduce a variational lower bound on the log-likelihood, which we optimize with stochastic gradient estimation. We demonstrate experimentally that our method can learn meaningful autoregressive orderings in image and graph generation. On the challenging domain of molecular graph generation, we achieve state-of-the-art results on the QM9 and ZINC250k benchmarks, evaluated across key metrics for distribution similarity and drug-likeless.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 64fa992f-beb7-478d-939e-215e334413e2Cited by top-tier papers10
- Any-Order Flexible Length Masked DiffusionJaeyeon Kim, Cheuk Lee Kit, Carles Domingo-Enrich, Yilun Du et al.ICLR 2026 · 51 citations
- Why Masking Diffusion Works: Condition on the Jump Schedule for Improved Discrete DiffusionAlan Nawzad Amin, Nate Gruver, Andrew Gordon WilsonNeurIPS 2025 · 18 citations
- Rainbow Padding: Mitigating Early Termination in Instruction-Tuned Diffusion LLMsBumjun Kim, Dongjae Jeon, Dueun Kim, Wonje Jeung et al.ICLR 2026 · 11 citations
- Forward-Learned Discrete Diffusion: Learning how to noise to denoise fasterGrigory Bartosh, Teodora Pandeva, Sushrut Karmalkar, Javier ZazoICLR 2026 · 4 citations
- Set Diffusion: Interpolating Token Orderings between Autoregression and Diffusion for Fast and Flexible DecodingMarianne Arriola, Volodymyr KuleshovICML 2026 · 2 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
- Simple and Effective Masked Diffusion Language ModelsSubham S. Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan et al.NeurIPS 2024 · 929 citations
- Simplified and Generalized Masked Diffusion for Discrete DataJiaxin Shi, Kehang Han, Zhe Wang, Arnaud Doucet et al.NeurIPS 2024 · 693 citations
Related papers
- Order Matters: Probabilistic Modeling of Node Sequence for Graph GenerationXiaohui Chen, Xu Han, Jiajing Hu, Francisco J. R. Ruiz et al.ICML 2021 · 40 citations
- Discovering Non-monotonic Autoregressive Orderings with Variational InferenceXuanlin Li, Brandon Trabucco, Dong Huk Park, Michael Luo et al.ICLR 2021 · 17 citations
- Graph Generative Pre-trained TransformerXiaohui Chen, Yinkai Wang, Jiaxing He, Yuanqi Du et al.ICML 2025
- Autoregressive Diffusion Model for Graph GenerationLingkai Kong, Jiaming Cui, Haotian Sun, Yuchen Zhuang et al.ICML 2023 · 105 citations
- GraphAF: a Flow-based Autoregressive Model for Molecular Graph GenerationChence Shi, Minkai Xu, Zhaocheng Zhu, Weinan Zhang et al.ICLR 2020 · 532 citations
