Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech
Xuanru Zhou, Jiachen Lian, Henry Hong, Xinyi Yang, Gopala Anumanchipalli
摘要
The man in the audio is speaking English and saying, "I know! Don't switch hands, okay?". It is difficult to determine the exact context of the conversation without further information. The audio clip seems to be a humorous exchange between two people, possibly a teacher and a student, as indicated by the phrase "don't switch hands, okay?" This could be a playful instruction in a classroom setting, where the teacher is trying to keep the student focused on a task. The audio is in a communication context that likely involves physical activity or interaction, such as passing objects, the speaker's fearful tone conveys a sense of urgency and insecurity. While on the surface their question, "don't switch hands okay?", simply asks for a 'yes' or 'no' answer, its underlying intent is to seek reassurance and a clear commitment from the other party. This implies a need for consistency and trust in the arrangement, as the speaker is looking for a definitive answer to resolve their fear and feel more secure in the situation. Speech World Model Ours It sounds like someone is giving instructions, possibly to someone else who's performing a task that requires focus, like a sport or a game. The speaker seems to be emphasizing not to switch hands, which might suggest they are guiding or coaching someone.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
相关 Paper
- Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni PerceptionZhen Wan, Chao-Han Huck Yang, Jinchuan Tian, Hanrong Ye 等ACL 2026 · 被引用 2 次
- Expressive Auditory Gestures in a Voice-Based Pedagogical AgentJessy Ceha, Edith LawCHI 2022 · 被引用 14 次
- How About Kind of Generating Hedges using End-to-End Neural Models?Alafate Abulimiti, Chloé Clavel, Justine CassellACL 2023 · 被引用 2 次
- Look Before You Speak: Visually Contextualized UtterancesPaul Hongsuck Seo, Arsha Nagrani, Cordelia SchmidCVPR 2021
- Speaker Information Can Guide Models to Better Inductive Biases: A Case Study On Predicting Code-SwitchingAlissa Ostapenko, Shuly Wintner, Melinda Fricke, Yulia TsvetkovACL 2022 · 被引用 6 次
