ICML2026
Text Generation as Continuous Latent Dynamics via Reinforcement Learning
Chen Jia
摘要
We propose to model text generation as a continuous-time latent dynamical process, where token generation is formulated as a Markov decision process whose internal state evolves via a neural ODE. This formulation bridges discrete token sequences and continuous semantic evolution, providing a theoretically grounded framework for text generation with continuous-time latent states. The framework is optimized via reinforcement learning, maximizing a composite objective that integrates cumulative rewards with a Kullback–Leibler divergence regularization term from a pre-trained language model. Both theoretical and empirical results demonstrate that our Continuous-Time Latent Language Model (CT-LLM) achieves superior effectiveness and efficiency in text generation, establishing a new paradigm for continuous-time language modeling.