Adaptive Thinking: Large Language Models Know When to Think in Latent Space
Pingzhi Li, Bairu Hou, Yun Zhu, Yihao Feng, Ke Ye, Tao Lei, Zhifeng Chen, Tianlong Chen, Xianzhi Du
摘要
Recent advances in large language models (LLMs) test-time computing have introduced the capability to perform intermediate chain-of-thought (CoT) reasoning (thinking) before generating answers. While increasing the thinking budget yields smooth performance improvements at inference time, the relationship between LLM capability, query complexity, and optimal budget allocation remains poorly understood for achieving compute-optimal inference. To address this challenge, we utilize , the agreement among multiple reasoning paths, as a proxy for thinking necessity. We first identify that lower self-consistency indicates when queries require extended thinking to reach correct answers. Building on this insight, we introduce (Self-Consistency-Guided Adapter for Thinking Allocation), a lightweight approach that adaptively allocates thinking budgets to optimize the performance-efficiency tradeoff. includes an adapter trained offline on a calibration dataset to predict self-consistency directly from the last layer hidden representations during the query prefilling stage. This prediction then guides on-the-fly budget allocation before thinking. The adapter is general, transferable across diverse tasks once trained, and introduces <1$$\textperthousand computational overhead during inference. Notably, Sonata is compatible with existing CoT compression methods, enabling further efficiency gains when managing thinking budgets across queries. Extensive experiments on multiple models (Qwen3-8B, Qwen3-32B, GPT-OSS-120B, Qwen3-235B-A22B) and benchmarks (AIME25, GSM8K, MATH500, GPQA, LiveCodeBench) demonstrate that achieves to reduction in thinking tokens while maintaining the same accuracy, or up to improvement in accuracy with the same token cost.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Deep Think with ConfidenceYichao Fu, Xuewei Wang, Hao Zhang, Yuandong Tian 等ICLR 2026 · 被引用 171 次
- Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept SpaceZhen Zhang, Xuehai He, Weixiang Yan, Ao Shen 等NeurIPS 2025 · 被引用 130 次
相关 Paper
- Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking TokensWei-Lin Chen, Liqian Peng, Tian Tan, Chao Zhao 等ICML 2026 · 被引用 20 次
- Plan and Budget: Effective and Efficient Test-Time Scaling on Reasoning Large Language ModelsJunhong Lin, Xinyue Zeng, Jie Zhu, Song Wang 等ICLR 2026 · 被引用 30 次
- Thought calibration: Efficient and confident test-time scalingMenghua Wu, Cai Zhou, Stephen Bates, Tommi S. JaakkolaEMNLP 2025 · 被引用 12 次
- RouteGoT: Node-Adaptive Routing for Cost-Efficient Graph of Thoughts ReasoningYuhang Liu, Ruijie Wang, Yunlong Chu, Bing Hao 等KDD 2026 · 被引用 1 次
- Let's Sample Step by Step: Adaptive-Consistency for Efficient Reasoning and Coding with LLMsPranjal Aggarwal, Aman Madaan, Yiming Yang, MausamEMNLP 2023 · 被引用 5 次
