Reinforced Context Order Recovery for Adaptive Reasoning and Planning
Long Ma, Fangwei Zhong, Yizhou Wang
Abstract
Modern causal language models, followed by rapid developments in discrete diffusion models, can now produce a wide variety of interesting and useful content. However, these families of models are predominantly trained to output tokens with a fixed (left-to-right) or random order, which may deviate from the logical order in which tokens are generated originally. In this paper, we observe that current causal and diffusion models encounter difficulties in problems that require adaptive token generation orders to solve tractably, which we characterize with the -information framework. Motivated by this, we propose Reinforced Context Order Recovery (ReCOR), a reinforcement-learning-based framework to extract adaptive, data-dependent token generation orders from text data without annotations. Self-supervised by token prediction statistics, ReCOR estimates the hardness of predicting every unfilled token and adaptively selects the next token during both training and inference. Experiments on challenging reasoning and planning datasets demonstrate the superior performance of ReCOR compared with baselines, sometimes outperforming oracle models supervised with the ground-truth order.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 07f97138-c3d9-4959-a158-a677caaf82eeCited by top-tier papers3
- Any-Order Flexible Length Masked DiffusionJaeyeon Kim, Cheuk Lee Kit, Carles Domingo-Enrich, Yilun Du et al.ICLR 2026 · 51 citations
- Rainbow Padding: Mitigating Early Termination in Instruction-Tuned Diffusion LLMsBumjun Kim, Dongjae Jeon, Dueun Kim, Wonje Jeung et al.ICLR 2026 · 11 citations
- Stop Training for the Worst: Progressive Unmasking Accelerates Masked Diffusion TrainingJaeyeon Kim, Jonathan Geuter, David Alvarez-Melis, Sham Kakade et al.ICML 2026 · 8 citations
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
Related papers
- Discovering Non-monotonic Autoregressive Orderings with Variational InferenceXuanlin Li, Brandon Trabucco, Dong Huk Park, Michael Luo et al.ICLR 2021 · 17 citations
- Beyond Fully Random Masking: Attention-Guided Denoising and Optimization for Diffusion Language ModelsJia Deng, Junyi Li, Xin Zhao, Jinpeng Wang et al.ACL 2026
- Diffusion Generative Recommendation with Continuous TokensHaohao Qu, Shanru Lin, Yujuan Ding, Yiqi Wang et al.WWW 2026 · 5 citations
- Unifying Masked Diffusion Models with Various Generation Orders and BeyondChunsan Hong, Sanghyun Lee, Jong Chul YEICML 2026
- Diffusion Forcing: Next-token Prediction Meets Full-Sequence DiffusionBoyuan Chen, Diego Marti Monso, Yilun Du, Max Simchowitz et al.NeurIPS 2024 · 751 citations
