The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning
Milad Aghajohari, Kamran Chitsaz, Amirhossein Kazemnejad, Sarath Chandar, Alessandro Sordoni, Aaron C. Courville, Siva Reddy
Abstract
Reinforcement learning (RL) has recently become a strong recipe for training reasoning LLMs that produce long chains of thought (LongCoT). Yet the standard RL "thinking environment", where the state is the prompt plus all prior reasoning tokens, makes the state unbounded and forces attention-based policies to pay quadratic compute as thoughts lengthen. We revisit the environment itself. We propose Markovian Thinking, a paradigm in which the policy advances reasoning while conditioning on a constant-size state, decoupling thinking length from context size. As an immediate consequence this yields linear compute with constant memory. We instantiate this idea with Delethink, an RL environment that structures reasoning into fixed-size chunks. Within each chunk, the model thinks as usual; at the boundary, the environment resets the context and reinitializes the prompt with a short carryover. Through RL, the policy learns to write a textual state near the end of each chunk sufficient for seamless continuation of reasoning after reset. Trained in this environment, an R1-Distill 1.5B model reasons in 8K-token chunks yet thinks up to 24K tokens, matching or surpassing LongCoT-RL trained with a 24K budget. With test-time scaling, Delethink continues to improve where LongCoT plateaus. The effect of linear compute is substantial: we empirically estimate at 96K average thinking length LongCoT-RL costs 27 H100-months vs. 7 for Delethink. Analysis at RL initialization shows off-the-shelf reasoning models (1.5B-120B) often sample Markovian traces zero-shot across diverse benchmarks, providing positive samples that make RL effective at scale. Our results show that redesigning the thinking environment is a powerful lever: it enables very long reasoning without quadratic overhead and opens a path toward efficient, scalable reasoning LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e7c8abc5-eaac-458b-8b4a-87967038c42bCited by top-tier papers9
- PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated ReasoningJingcheng Hu, Yinmin Zhang, Shijie Shang, Xiaobo Yang et al.ACL 2026 · 15 citations
- Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RLIan Wu, Yuxiao Qu, Amrith Setlur, Aviral KumarICML 2026 · 8 citations
- LongCoT: Benchmarking Long-Horizon Chain-of-Thought ReasoningSumeet Motwani, Daniel Nichols, Charles London, Peggy Li et al.ICML 2026 · 2 citations
- Learning to Correct: Reinforcement Learning for Multi-Attempt Chain-of-ThoughtMuhammed Emrullah Ildiz, Halil Alperen Gozeten, Ege Onur Taga, Samet OymakICML 2026 · 2 citations
- The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-ThoughtMoritz Brösamle, Stephan EcksteinICML 2026
Builds on18
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
Related papers
- Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RLSongjun Tu, Jiahao Lin, Qichao Zhang, Xiangyu Tian et al.NeurIPS 2025 · 69 citations
- LEASH: Adaptive Length Penalty and Reward Shaping for Efficient Large Reasoning ModelYanhao Li, Lu Ma, Jiaran Zhang, Lexiang Tang et al.ACL 2026 · 8 citations
- Long-Context Reasoning Through Proxy-Based Chain-of-Thought TuningMiao Li, Irina Saparina, Alexander Gurung, Mirella LapataACL 2026
- When More is Less: Understanding Chain-of-Thought Length in LLMsYuyang Wu, Yifei Wang, Ziyu Ye, Tianqi Du et al.ICLR 2026 · 225 citations
- InftyThink: Breaking the Length Limits of Long-Context Reasoning in Large Language ModelsYuchen Yan, Yongliang Shen, Yang Liu, Jin Jiang et al.ICLR 2026 · 48 citations
