Lune

ICML2026Top-tier venue

Beyond Test-Time Memory: State-Space Optimal Control for LLM Reasoning

Peihao Wang, Shan Yang, Xijun Wang, Tesi Xiao, Jiahui Gao, Changlong Yu, Yu Lou, Pan Li, Zhangyang “Atlas” Wang, Ming Lin, Rene Vidal

2026Year

Abstract

Associative memory has long underpinned the design of sequential models. Beyond recall, humans reason by projecting future states and selecting goal-directed actions, a capability that modern language models increasingly require but do not natively encode. While prior work uses reinforcement learning or test-time training, planning remains external to the model architecture. We formulate reasoning as optimal control and introduce the Test-Time Control (TTC) layer, which performs finite-horizon LQR planning over latent states at inference time, represents a value function within neural architectures, and leverages it as the nested objective to enable planning before prediction. To ensure scalability, we derive a hardware-efficient LQR solver based on a symplectic formulation and implement it as a fused CUDA kernel, enabling parallel execution with minimal overhead. Integrated as an adapter into pretrained LLMs, TTC layers improve mathematical reasoning performance by up to +27.8% on MATH-500 and 2-3× Pass@8 improvements on AMC and AIME, demonstrating that embedding optimal control as an architectural component provides an effective and scalable mechanism for reasoning beyond test-time training. This memory-centric paradigm has proven highly effective for language modeling, yet increasingly reveals limitations when models are required to reason, discover, or solve problems. These tasks demand mechanisms beyond memorization and retrieval. From a cognitive perspective, human intelligence operates through an interplay between System 1 and System 2 thinking (Kahneman, 2011) . System 1 relies on fast, automatic pattern matching over memory, while System 2 engages in deliberate, multi-step planning and long-horizon reasoning. Current LLM architectures largely instantiate System 1 behavior: they predict the next token by extrapolating from past context, but lack a dedicated architectural mechanism for System 2-style planning.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 27c36a5c-fd35-4fd4-a4ed-1fb528804822

Builds on37

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines