R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Token Routing
Tianyu Fu, Yi Ge, Yichen You, Enshu Liu, Zhihang Yuan, Guohao Dai, Shengen Yan, Huazhong Yang, Yu Wang
Abstract
Large Language Models (LLMs) achieve impressive reasoning capabilities at the cost of substantial inference overhead, posing substantial deployment challenges. Although distilled Small Language Models (SLMs) significantly enhance efficiency, their performance suffers as they fail to follow LLMs'reasoning paths. Luckily, we reveal that only a small fraction of tokens genuinely diverge reasoning paths between LLMs and SLMs. Most generated tokens are either identical or exhibit neutral differences, such as minor variations in abbreviations or expressions. Leveraging this insight, we introduce Roads to Rome (R2R), a neural token routing method that selectively utilizes LLMs only for these critical, path-divergent tokens, while leaving the majority of token generation to the SLM. We also develop an automatic data generation pipeline that identifies divergent tokens and generates token-level routing labels to train the lightweight router. We apply R2R to combine R1-1.5B and R1-32B models from the DeepSeek family, and evaluate on challenging math, coding, and QA benchmarks. With an average activated parameter size of 5.6B, R2R surpasses the average accuracy of R1-7B by 1.6x, outperforming even the R1-14B model. Compared to R1-32B, it delivers a 2.8x wall-clock speedup with comparable performance, advancing the Pareto frontier of test-time scaling efficiency. Our code is available at https://github.com/thu-nics/R2R.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 00852273-c3f5-4459-9172-0446042265a2Cited by top-tier papers9
- Cache-to-Cache: Direct Semantic Communication Between Large Language ModelsTianyu Fu, Zihan Min, Hanling Zhang, Jichao Yan et al.ICLR 2026 · 59 citations
- Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language ModelsTianyu Fu, Yichen You, Zekai Chen, Guohao Dai et al.ICML 2026 · 20 citations
- PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation ModelsTianchen Zhao, Ke Hong, Xinhao Yang, Xuefeng Xiao et al.NeurIPS 2025 · 19 citations
- FROST: Filtering Reasoning Outliers with Attention for Efficient ReasoningHaozheng Luo, Zhuolin Jiang, Md Zahid Hasan, Yan Chen et al.ICLR 2026 · 4 citations
- Dissecting Failure Dynamics in Large Language Model ReasoningWei Zhu, Jian Zhang, Lixing Yu, Kun Yue et al.ACL 2026 · 2 citations
Builds on21
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningShenzhi Wang, Le Yu, Chang Gao, Chujie Zheng et al.NeurIPS 2025 · 592 citations
Related papers
- TokenSqueeze: Performance-Preserving Compression for Reasoning LLMsYuxiang Zhang, Zhengxu Yu, Weihang Pan, Zhongming Jin et al.NeurIPS 2025 · 6 citations
- Select to Think: Unlocking SLM Potential with Local SufficiencyWenxuan Ye, Yangyang Zhang, Xueli An, Georg Carle et al.ICML 2026 · 2 citations
- Think When Needed: Model-Aware Reasoning Routing for LLM-based RankingHuizhong Guo, Tianjun Wei, Dongxia Wang, Yingpeng Du et al.SIGIR 2026
- RouteLLM: Learning to Route LLMs from Preference DataIsaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang et al.ICLR 2025
- Let the LLM Stick to Its Strengths: Learning to Route Economical LLMYi-Kai Zhang, Shiyin Lu, Qingguo Chen, Weihua Luo et al.NeurIPS 2025 · 3 citations
