Lune

ICLR2026Top-tier venue

Learning to Reason over Continuous Tokens with Reinforcement Learning

Yiran Zhao, Yuhui Xu, Doyen Sahoo, Caiming Xiong, Junnan Li

2026Year
1Citations

Abstract

Large Language Models (LLMs) have shown strong performance in complex reasoning tasks, especially when guided by Chain-of-Thought (CoT) prompting. However, conventional CoT reasoning in the discrete token space suffers from high computational and memory costs due to verbose intermediate steps. Recent work has explored latent reasoning in the embedding space to improve efficiency, but often at the cost of clarity and performance. In this work, we propose Hybrid Reasoning (HyRea), a unified framework that enables LLMs to dynamically switch between explicit (token-based) and latent (embedding-based) reasoning during inference. To train the model to make these decisions effectively, we introduce a two-stage training pipeline: (1) a supervised cold-start phase that introduces latent reasoning by replacing low-entropy CoT steps with embeddings, and (2) a reinforcement learning phase using Group Relative Policy Optimization (GRPO) to fine-tune the model's reasoning strategy based on task-specific rewards. Experiments on mathematical reasoning benchmarks show that HyRea achieves significant reductions in token usage while maintaining or improving accuracy, offering an effective and scalable solution for efficient multi-step reasoning in LLMs. Our code is publicly available at https://github.com/zhaoyiran924/HyRea.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 49085119-36a2-4757-b975-85f152c55a0a

Builds on18

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines