Shanks: Simultaneous Hearing and Thinking for Spoken Language Models
Cheng-Han Chiang, Xiaofei Wang, Linjie Li, Chung-Ching Lin, Kevin Lin, Shujie Liu, Zhendong Wang, Zhengyuan Yang, Hung-yi Lee, Lijuan Wang
摘要
Current large language models (LLMs) and spoken language models (SLMs) begin thinking and taking actions only after the user has finished their turn. This disables the model from interacting with the user during the user's turn and can lead to a high response latency for waiting for the model to think. Consequently, thinking after receiving the full input is not suitable for speech-to-speech interaction, where real-time and low-latency interaction is important. We address the above issue by drawing inspiration from the fact that humans can naturally "think while listening". In this paper, we propose SHANKS, a general inference framework that enables SLMs to generate unspoken chain-of-thought reasoning when listening to the user input. SHANKS streams the input speech in fixed-duration chunks and, as soon as a chunk is received, generates unspoken reasoning based on all previous speech and reasoning, while the user continues speaking. SHANKS uses unspoken reasoning to determine whether to interrupt the user and make tool calls to complete the task. We demonstrate that SHANKS enhances the real-time user-SLM interaction in two scenarios: (1) When the user is presenting their step-by-step solution to a math problem, SHANKS can listen to and reason over the user's speech and make an interruption when the user makes a mistake. SHANKS interrupts the user 37.1% more accurately compared with a baseline that interrupts the user without thinking. (2) In a tool-augmented dialogue scenario, where the model needs to make tool calls to achieve the user's request, SHANKS can complete 56.9% of the tool calls before the user even ends their turn. Overall, SHANKS is a step toward models that keep thinking throughout the conversation, not only after a turn ends. Animated illustrations of SHANKS can be found at https: //d223302.github.io/SHANKS/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
相关 Paper
- STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language ModelsCheng-Han Chiang, Xiaofei Wang, Linjie Li, Chung-Ching Lin 等ICLR 2026 · 被引用 38 次
- Can Speech LLMs Think while Listening?Yi-Jen Shih, Desh Raj, Chunyang Wu, Wei Zhou 等ICLR 2026 · 被引用 24 次
- Think, Verbalize, then Speak: Bridging Complex Thoughts and Comprehensible SpeechTony Woo, Sehun Lee, Kang-Wook Kim, Gunhee KimEMNLP 2025
- PLANTAIN: Plan-Answer Interleaved ReasoningAnthony Liang, Jonathan Berant, Adam Fisch, Abhimanyu Goyal 等ICML 2026 · 被引用 1 次
- StreamingThinker: Large Language Models Can Think While ReadingJunlong Tong, Yingqi Fan, Anhao Zhao, Yunpu Ma 等ICLR 2026 · 被引用 17 次
