Shanks: Simultaneous Hearing and Thinking for Spoken Language Models
Cheng-Han Chiang, Xiaofei Wang, Linjie Li, Chung-Ching Lin, Kevin Lin, Shujie Liu, Zhendong Wang, Zhengyuan Yang, Hung-yi Lee, Lijuan Wang
Abstract
Current large language models (LLMs) and spoken language models (SLMs) begin thinking and taking actions only after the user has finished their turn. This disables the model from interacting with the user during the user's turn and can lead to a high response latency for waiting for the model to think. Consequently, thinking after receiving the full input is not suitable for speech-to-speech interaction, where real-time and low-latency interaction is important. We address the above issue by drawing inspiration from the fact that humans can naturally "think while listening". In this paper, we propose SHANKS, a general inference framework that enables SLMs to generate unspoken chain-of-thought reasoning when listening to the user input. SHANKS streams the input speech in fixed-duration chunks and, as soon as a chunk is received, generates unspoken reasoning based on all previous speech and reasoning, while the user continues speaking. SHANKS uses unspoken reasoning to determine whether to interrupt the user and make tool calls to complete the task. We demonstrate that SHANKS enhances the real-time user-SLM interaction in two scenarios: (1) When the user is presenting their step-by-step solution to a math problem, SHANKS can listen to and reason over the user's speech and make an interruption when the user makes a mistake. SHANKS interrupts the user 37.1% more accurately compared with a baseline that interrupts the user without thinking. (2) In a tool-augmented dialogue scenario, where the model needs to make tool calls to achieve the user's request, SHANKS can complete 56.9% of the tool calls before the user even ends their turn. Overall, SHANKS is a step toward models that keep thinking throughout the conversation, not only after a turn ends. Animated illustrations of SHANKS can be found at https: //d223302.github.io/SHANKS/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
Related papers
- STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language ModelsCheng-Han Chiang, Xiaofei Wang, Linjie Li, Chung-Ching Lin et al.ICLR 2026 · 38 citations
- Can Speech LLMs Think while Listening?Yi-Jen Shih, Desh Raj, Chunyang Wu, Wei Zhou et al.ICLR 2026 · 24 citations
- Think, Verbalize, then Speak: Bridging Complex Thoughts and Comprehensible SpeechTony Woo, Sehun Lee, Kang-Wook Kim, Gunhee KimEMNLP 2025
- PLANTAIN: Plan-Answer Interleaved ReasoningAnthony Liang, Jonathan Berant, Adam Fisch, Abhimanyu Goyal et al.ICML 2026 · 1 citation
- StreamingThinker: Large Language Models Can Think While ReadingJunlong Tong, Yingqi Fan, Anhao Zhao, Yunpu Ma et al.ICLR 2026 · 17 citations
