Reasoning Palette: Modulating Reasoning via Latent Contextualization for Controllable Exploration for (V)LMs
Rujiao Long, Yang Li, Xingyao Zhang, Weixun Wang, Tianqianjin Lin, Xi Zhao, Yuchi Xu, Wenbo Su, Junchi Yan, Bo Zheng
Abstract
Exploration capacity shapes both inference-time performance and reinforcement learning (RL) training for large (vision-) language models, as stochastic sampling often yields redundant reasoning paths with little high-level diversity. This paper proposes Reasoning Palette, a novel latent-modulation framework that endows the model with a stochastic latent variable for strategic contextualization, guiding its internal planning prior to token generation. This latent context is inferred from the mean-pooled embedding of a question-answer pair via a variational autoencoder (VAE), where each sampled latent potentially encodes a distinct reasoning context. During inference, a sampled latent is decoded into learnable token prefixes and prepended to the input prompt, modulating the model's internal reasoning trajectory. In this way, the model performs internal sampling over reasoning strategies prior to output generation, which shapes the style and structure of the entire response sequence. A brief supervised fine-tuning (SFT) warm-up phase allows the model to adapt to this latent conditioning. Within RL optimization, Reasoning Palette facilitates structured exploration by enabling on-demand injection for diverse reasoning modes, significantly enhancing exploration efficiency and sustained learning capability. Experiments across multiple reasoning benchmarks demonstrate that our method enables interpretable and controllable control over the (vision-) language model's strategic behavior, thereby achieving consistent performance gains over standard RL methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Attention Illuminates LLM Reasoning: The Uncovered Preplan-and-Anchor Rhythm Enables Fine-Grained Policy OptimizationYang Li, Zhichen Dong, Yuhan Sun, Weixun Wang et al.ICML 2026 · 25 citations
- How Does Reasoning Flow? Tracing Attention-Induced Information Flow for Targeted RL in LLMsZhichen Dong, Yang Li, Yuhan Sun, Weixun Wang et al.ICML 2026 · 1 citation
- MaskCO: Masked Generation Drives Effective Representation Learning and Exploiting for Combinatorial OptimizationLvda Chen, Yang Li, Junchi YanICLR 2026
- Let's Explore Step by Step: Generating Provable Formal Statements with Deductive ExplorationQi Liu, Kangjie Bao, Yue Yang, Xinhao Zheng et al.ICLR 2026
Builds on31
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
Related papers
- Unveiling the Compositional Ability Gap in Vision-Language Reasoning ModelTianle Li, Jihai Zhang, Yongming Rao, Yu ChengNeurIPS 2025 · 17 citations
- VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool UseMingyuan Wu, Jingcheng Yang, Jize Jiang, Meitang Li et al.ICLR 2026 · 87 citations
- Learning Diverse Responses with Prefix-Conditioned Supervised Fine-TuningZhiyuan Fan, Guanqiao Chen, Yanyi Huang, Mingkuan Zhao et al.ACL 2026
- Hybrid Latent Reasoning via Reinforcement LearningZhenrui Yue, Bowen Jin, Huimin Zeng, Honglei Zhuang et al.NeurIPS 2025 · 28 citations
- Learning Structured Reasoning via Tractable Trajectory ControlPo-Nien Kung, Zhen Yang, Jeffrey Luo, Cheng-Fu Yang et al.ICML 2026
