Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech
Xuanru Zhou, Jiachen Lian, Henry Hong, Xinyi Yang, Gopala Anumanchipalli
Abstract
The man in the audio is speaking English and saying, "I know! Don't switch hands, okay?". It is difficult to determine the exact context of the conversation without further information. The audio clip seems to be a humorous exchange between two people, possibly a teacher and a student, as indicated by the phrase "don't switch hands, okay?" This could be a playful instruction in a classroom setting, where the teacher is trying to keep the student focused on a task. The audio is in a communication context that likely involves physical activity or interaction, such as passing objects, the speaker's fearful tone conveys a sense of urgency and insecurity. While on the surface their question, "don't switch hands okay?", simply asks for a 'yes' or 'no' answer, its underlying intent is to seek reassurance and a clear commitment from the other party. This implies a need for consistency and trust in the arrangement, as the speaker is looking for a definitive answer to resolve their fear and feel more secure in the situation. Speech World Model Ours It sounds like someone is giving instructions, possibly to someone else who's performing a task that requires focus, like a sport or a game. The speaker seems to be emphasizing not to switch hands, which might suggest they are guiding or coaching someone.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
Related papers
- Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni PerceptionZhen Wan, Chao-Han Huck Yang, Jinchuan Tian, Hanrong Ye et al.ACL 2026 · 2 citations
- Expressive Auditory Gestures in a Voice-Based Pedagogical AgentJessy Ceha, Edith LawCHI 2022 · 14 citations
- How About Kind of Generating Hedges using End-to-End Neural Models?Alafate Abulimiti, Chloé Clavel, Justine CassellACL 2023 · 2 citations
- Look Before You Speak: Visually Contextualized UtterancesPaul Hongsuck Seo, Arsha Nagrani, Cordelia SchmidCVPR 2021
- Speaker Information Can Guide Models to Better Inductive Biases: A Case Study On Predicting Code-SwitchingAlissa Ostapenko, Shuly Wintner, Melinda Fricke, Yulia TsvetkovACL 2022 · 6 citations
