DREAM-R: Multimodal Speculative Reasoning with RL-Based Refined Drafting, Precise Verification, and Fully Parallel Execution
Yunhai Hu, Zining Liu, Xiangyang Yin, Tianhua Xia, BO BAO, Eric Sather, Vithursan Thangarasa, Sai Qian Zhang
Abstract
Speculative reasoning has recently been proposed as a means to accelerate reasoning-intensive generation in large multimodal models, but its effectiveness is often constrained by misalignment between speculative drafts and target-verified reasoning. In this work, we introduce DREAM-R, a framework that substantially improves the performance of speculative reasoning. At its core, DREAM-R employs Speculative Alignment Policy Optimization (SAPO), a reinforcementlearning objective that trains draft models to generate reasoning steps that are both faithful to target trajectories and concise. We further propose a Contrastive Probability Normalization (CPN) that uses a ratio-based criterion to provide stable and interpretable acceptance of speculative steps only when positive evidence clearly dominates, thereby preventing error propagation. Building on these components, we develop a Fully Parallel Speculative Reasoning (FPSR) framework that parallelizes draft generation, target-side reasoning, and verification across multi-step reasoning, enabling early stopping and clean fallback. Experiments on reasoning-heavy benchmarks demonstrate up to 2.48× speedup while preserving target-model accuracy, yielding substantial efficiency gains without compromising reasoning quality. Code is available at https://github. com/HuYunhai-Alex/DREAM-R .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
Related papers
- Segment-Aligned Policy Optimization for Multi-Modal ReasoningLei Gao, Zhuoming Li, Mengxi Jia, Jiakang Yuan et al.ICML 2026 · 2 citations
- FastGRPO: Accelerating Policy Optimization via Concurrency-aware Speculative Decoding and Online Draft LearningYizhou Zhang, Ning Lv, Teng Wang, Jisheng DangICLR 2026 · 10 citations
- CFPO: Counterfactual Policy Optimization for Multimodal ReasoningZhangyuan Yu, Wanran Sun, Guangjing Yang, Xiaohu Wu et al.ICML 2026
- Slow-Fast Policy Optimization: Reposition-Before-Update for LLM ReasoningZiyan Wang, Zheng Wang, Xingwei Qu, Qi Cheng et al.ICLR 2026 · 4 citations
- Re-SpS: A Reinforcement Learning Approach to Speculative SamplingChenan Wang, Daniel H. Shi, Haipeng ChenAAAI 2026
