SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMs
Dachuan Shi, Abedelkadir Asi, Keying Li, Xiangchi Yuan, Leyan Pan, Wenke Lee, Wen Xiao
摘要
Recent work shows that, beyond discrete reasoning through explicit chain-of-thought steps, which are limited by the boundaries of natural languages, large language models (LLMs) can also reason continuously in latent space, allowing richer information per step and thereby improving token efficiency. Despite this promise, latent reasoning still faces two challenges, especially in training-free settings: 1) purely latent reasoning broadens the search distribution by maintaining multiple implicit paths, which diffuses probability mass, introduces noise, and impedes convergence to a single high-confidence solution, thereby hurting accuracy; and 2) overthinking persists even without explicit text, wasting tokens and degrading efficiency. To address these issues, we introduce SwiReasoning, a training-free framework for LLM reasoning which features two key innovations: 1) SwiReasoning dynamically switches between explicit and latent reasoning, guided by block-wise confidence estimated from entropy trends in next-token distributions, to balance exploration and exploitation and promote timely convergence. 2) By limiting the maximum number of thinking-block switches, SwiReasoning curbs overthinking and improves token efficiency across varying problem difficulties. On widely used mathematics, STEM, coding, and general benchmarks, SwiReasoning consistently improves average accuracy by 1.8%–3.1% across reasoning LLMs of different model families and scales. Furthermore, under constrained budgets, SwiReasoning improves average token efficiency by 57%-79%, with larger gains as budgets tighten.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware DecodingZhongxing Xu, Zhonghua Wang, Zhe Qian, Dachuan Shi 等CVPR 2026 · 被引用 16 次
- Latent Thoughts Tuning: Bridging Context and Reasoning with Fused Information in Latent TokensWeihao Liu, Dehai Min, Lu ChengICML 2026 · 被引用 3 次
- Restoring Exploration after Post-Training: Latent Exploration Decoding for Large Reasoning ModelsWenhui Tan, Fiorenzo Parascandolo, Enver Sangineto, Jianzhong Ju 等ICML 2026 · 被引用 2 次
- SeLaR: Selective Latent Reasoning in Large Language ModelsRenyu Fu, Guibo LuoACL 2026 · 被引用 2 次
它引用的顶会 Paper31
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
相关 Paper
- Beyond Tokens: Dynamic Latent Reasoning via Semantic Residual RefinementFangrui Lv, Lei Wang, Ruixin Hong, Yong Du 等AAAI 2026
- DyCon: Dynamic Reasoning Control via Evolving Difficulty ModelingTengyao Tu, Yulin Li, Huiling Zhen, Libo Qin 等ICML 2026
- Learning to Reason over Continuous Tokens with Reinforcement LearningYiran Zhao, Yuhui Xu, Doyen Sahoo, Caiming Xiong 等ICLR 2026 · 被引用 1 次
- SABER: Switchable and Balanced Training for Efficient LLM ReasoningKai Zhao, Yanjun Zhao, Jiaming Song, Shien He 等AAAI 2026 · 被引用 9 次
- Efficient Reasoning with Balanced ThinkingYulin Li, Tengyao Tu, Li Ding, Junjie Wang 等ICLR 2026 · 被引用 7 次
