Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive Flows
Ruixiang Zhang, Shuangfei Zhai, Jiatao Gu, Yizhe Zhang, Huangjie Zheng, Tianrong Chen, Miguel Ángel Bautista, Joshua Susskind, Navdeep Jaitly
Abstract
Autoregressive models have driven remarkable progress in language modeling. Their foundational reliance on discrete tokens, unidirectional context, and singlepass decoding, while central to their success, also inspires the exploration of a design space that could offer new axes of modeling flexibility. In this work, we explore an alternative paradigm, shifting language modeling from a discrete token space to a continuous latent space. We propose a novel framework TarFlowLM, that employs transformer-based autoregressive normalizing flows [71] to model these continuous representations. This approach unlocks substantial flexibility, enabling the construction of models that can capture global bi-directional context through stacked, alternating-direction autoregressive transformations, support block-wise generation with flexible token patch sizes, and facilitate a hierarchical multi-pass generation process. We further propose new mixture-based coupling transformations designed to capture complex dependencies within the latent space shaped by discrete data, and demonstrate theoretical connections to conventional discrete autoregressive models. Extensive experiments on language modeling benchmarks demonstrate strong likelihood performance and highlight the flexible modeling capabilities inherent in our framework.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Continuously Augmented Discrete Diffusion model for Categorical Generative ModelingHuangjie Zheng, Shansan Gong, Ruixiang Zhang, Tianrong Chen et al.ICLR 2026 · 30 citations
- FARMER: Flow AutoRegressive Transformer over PixelsGuangting Zheng, Qinyu Zhao, Tao Yang, Fei Xiao et al.CVPR 2026 · 17 citations
- Enhancing Reasoning for Diffusion LLMs via Distribution Matching Policy OptimizationYuchen Zhu, Wei Guo, Jaemoo Choi, Petr Molodyk et al.ICML 2026 · 13 citations
- Normalizing Flows with Iterative DenoisingTianrong Chen, Jiatao Gu, David Berthelot, Joshua M Susskind et al.ICML 2026 · 3 citations
Builds on37
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang et al.NeurIPS 2022 · 1,546 citations
Related papers
- Normalizing Flows are Capable Generative ModelsShuangfei Zhai, Ruixiang Zhang, Preetum Nakkiran, David Berthelot et al.ICML 2025
- STARFlow: Scaling Latent Normalizing Flows for High-resolution Image SynthesisJiatao Gu, Tianrong Chen, David Berthelot, Huangjie Zheng et al.NeurIPS 2025 · 34 citations
- Bidirectional Normalizing Flow: From Data to Noise and BackYiyang Lu, Qiao Sun, Xianbang Wang, Zhicheng Jiang et al.CVPR 2026 · 7 citations
- KALL-E: Autoregressive Speech Synthesis with Next-Distribution PredictionKangxiang Xia, Xinfa Zhu, Jixun Yao, Wenjie Tian et al.AAAI 2026 · 3 citations
- Optimality of FSQ Tokens for Continuous Diffusion for Categorical Data with Application to Text-to-SpeechVadim Popov, Wenju Gu, Tasnima Sadekova, Georgii Aparin et al.ICML 2026
