Recent Advances in Speech Language Models: A Survey
Wenqian Cui, Dianzhi Yu, Xiaoqi Jiao, Ziqiao Meng, Guangyan Zhang, Qichao Wang, Steven Y. Guo, Irwin King
摘要
Large Language Models (LLMs) have recently garnered significant attention, primarily for their capabilities in textbased interactions. However, natural human interaction often relies on speech, necessitating a shift towards voice-based models. A straightforward approach to achieve this involves a pipeline of "Automatic Speech Recognition (ASR) + LLM + Text-to-Speech (TTS)", where input speech is transcribed to text, processed by an LLM, and then converted back to speech. Despite being straightforward, this method suffers from inherent limitations, such as information loss during modality conversion, significant latency due to the complex pipeline, and error accumulation across the three stages. To address these issues, Speech Language Models (SpeechLMs)-end-to-end models that generate speech without converting from text-have emerged as a promising alternative. This survey paper provides the first comprehensive overview of recent methodologies for constructing SpeechLMs, detailing the key components of their architecture and the various training recipes integral to their development. Additionally, we systematically survey the various capabilities of SpeechLMs, categorize their evaluation metrics, and discuss the challenges and future research directions in this rapidly evolving field.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning BenchmarkDingdong Wang, Junan Li, Jincenzi Wu, Dongchao Yang 等ICLR 2026 · 被引用 143 次
- Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video GamesDongmin Park, Minkyu Kim, Beongjun Choi, Junhyuck Kim 等ICLR 2026 · 被引用 30 次
- Can Speech LLMs Think while Listening?Yi-Jen Shih, Desh Raj, Chunyang Wu, Wei Zhou 等ICLR 2026 · 被引用 24 次
- EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language ModelsLi Zhou, Lutong Yu, You Lyu, Yihang Lin 等ICLR 2026 · 被引用 13 次
- Closing the Modality Reasoning Gap for Speech Large Language ModelsChaoren Wang, Heng Lu, Xueyao Zhang, Shujie Liu 等ACL 2026 · 被引用 10 次
它引用的顶会 Paper44
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao 等ICLR 2021 · 被引用 1,902 次
相关 Paper
- Towards True Speech-to-Speech Models Without Text GuidanceXingjian Zhao, Zhe Xu, Luozhijie Jin, Yang Wang 等ICLR 2026 · 被引用 8 次
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech InteractionChaoyou Fu, Haojia Lin, Xiong Wang, Yifan Zhang 等NeurIPS 2025 · 被引用 234 次
- LLaMA-Omni: Seamless Speech Interaction with Large Language ModelsQingkai Fang, Shoutao Guo, Yan Zhou, Zhengrui Ma 等ICLR 2025 · 被引用 2 次
- Distilling an End-to-End Voice Assistant Without Instruction Training DataWilliam Barr Held, Yanzhe Zhang, Weiyan Shi, Minzhi Li 等ACL 2025
- Speech Recognition Meets Large Language Model: Benchmarking, Models, and ExplorationZiyang Ma, Guanrou Yang, Yifan Yang, Zhifu Gao 等AAAI 2025 · 被引用 17 次
