SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation
Sirry Chen, Jieyi Wang, Wei Chen, Zhongyu Wei
摘要
Medical consultations are intrinsically speechcentric. However, most prior works focus on long-text-based interactions, which are cumbersome and patient-unfriendly. Recent advances in speech language models (SpeechLMs) have enabled more natural speech-based interaction, yet the scarcity of medical speech data and the inefficiency of directly fine-tuning on speech data jointly hinder the adoption of SpeechLMs in medical consultation. In this paper, we propose SpeechMedAssist, a SpeechLM natively capable of conducting speech-based multi-turn interactions with patients. By exploiting the architectural properties of SpeechLMs, we decouple the conventional one-stage training into a two-stage paradigm consisting of (1) Knowledge & Capability Injection via Text and (2) Modality Re-alignment with Limited Speech Data, thereby reducing the requirement for medical speech data to only 10k synthesized samples. To evaluate SpeechLMs for medical consultation scenarios, we design a benchmark comprising both single-turn question answering and multi-turn simulated interactions. Experimental results show that our model outperforms all baselines in both effectiveness and robustness in most evaluation settings. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper2
相关 Paper
- Ask Patients with Patience: Enabling LLMs for Human-Centric Medical Dialogue with Grounded ReasoningJiayuan Zhu, Jiazhen Pan, Yuyuan Liu, Fenglin Liu 等EMNLP 2025 · 被引用 1 次
- LLaMA-Omni 2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech SynthesisQingkai Fang, Yan Zhou, Shoutao Guo, Shaolei Zhang 等ACL 2025
- Scaling Speech-Text Pre-training with Synthetic Interleaved DataAohan Zeng, Zhengxiao Du, Mingdao Liu, Lei Zhang 等ICLR 2025
- Note2Chat: Improving LLMs for Multi-Turn Clinical History Taking Using Medical NotesYang Zhou, Zhenting Sheng, Mingrui Tan, Yuting Song 等AAAI 2026
- Unlocking Speech–Text Compositional Powers: Instruction-Following Speech Language Models without Instruction TuningCongrui Du, Yang Zhang, Kaizhi Qian, Shiyu ChangICML 2026
