Text Generation as Continuous Latent Dynamics via Reinforcement Learning
Chen Jia
Abstract
We propose to model text generation as a continuous-time latent dynamical process, where token generation is formulated as a Markov decision process whose internal state evolves via a neural ODE. This formulation bridges discrete token sequences and continuous semantic evolution, providing a theoretically grounded framework for text generation with continuous-time latent states. The framework is optimized via reinforcement learning, maximizing a composite objective that integrates cumulative rewards with a Kullback–Leibler divergence regularization term from a pre-trained language model. Both theoretical and empirical results demonstrate that our Continuous-Time Latent Language Model (CT-LLM) achieves superior effectiveness and efficiency in text generation, establishing a new paradigm for continuous-time language modeling.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on18
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang et al.NeurIPS 2022 · 1,546 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
Related papers
- Unifying Continuous and Discrete Text Diffusion with Non-simultaneous Diffusion ProcessesBocheng Li, Zhujin Gao, Linli XuACL 2025
- MotionStreamer: Streaming Motion Generation via Diffusion-Based Autoregressive Model in Causal Latent SpaceLixing Xiao, Shunlin Lu, Huaijin Pi, Ke Fan et al.ICCV 2025 · 11 citations
- Flowing Through States: Neural ODE Regularization for Reinforcement LearningMohamed Ghanem, Bernd FinkbeinerICLR 2026
- Optimality of FSQ Tokens for Continuous Diffusion for Categorical Data with Application to Text-to-SpeechVadim Popov, Wenju Gu, Tasnima Sadekova, Georgii Aparin et al.ICML 2026
- KALL-E: Autoregressive Speech Synthesis with Next-Distribution PredictionKangxiang Xia, Xinfa Zhu, Jixun Yao, Wenjie Tian et al.AAAI 2026 · 3 citations
