Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
Dingdong Wang, Junan Li, Mingyu Cui, Dongchao Yang, Xueyuan Chen, Helen M. Meng
Abstract
With the rise of Speech Large Language Models (SpeechLLMs), two dominant approaches have emerged for speech processing: discrete tokens and continuous features. Each approach has demonstrated strong capabilities in audiorelated processing tasks. However, the performance gap between these two paradigms has not been thoroughly explored. To address this gap, we present a fair comparison of selfsupervised learning (SSL)-based discrete and continuous features under the same experimental settings. We evaluate their performance across six spoken language understandingrelated tasks using both small and large-scale LLMs (Qwen1.5-0.5B and Llama3.1-8B). We further conduct in-depth analyses, including efficient comparison, SSL layer analysis, LLM layer analysis, and robustness comparison. Our findings reveal that continuous features generally outperform discrete tokens in various tasks. Each speech processing method exhibits distinct characteristics and patterns in how it learns and processes speech information. We hope our findings will provide valuable insights to advance spoken language understanding in SpeechLLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9e8ed2a3-76f2-4b58-bd86-3e0494efb903Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- SALMONN: Towards Generic Hearing Abilities for Large Language ModelsChangli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen et al.ICLR 2024 · 557 citations
- InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-trainingDingdong Wang, Jin Xu, Ruihang Chu, Zhifang Guo et al.ACL 2025 · 9 citations
- Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language ModelsSander Land, Max BartoloEMNLP 2024 · 4 citations
- ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language ModelingDongchao Yang, Songxiang Liu, Haohan Guo, Jiankun Zhao et al.ICML 2025
Related papers
- A Variational Framework for Improving Naturalness in Generative Spoken Language ModelsLi-Wei Chen, Takuya Higuchi, Zakaria Aldeneh, Ahmed Hussen Abdelaziz et al.ICML 2025
- Generative Spoken Language Model based on continuous word-sized audio tokensRobin Algayres, Yossi Adi, Tu Anh Nguyen, Jade Copet et al.EMNLP 2023 · 3 citations
- SyllableLM: Learning Coarse Semantic Units for Speech Language ModelsAlan Baade, Puyuan Peng, David HarwathICLR 2025
- Layer-wise Minimal Pair Probing Reveals Contextual Grammatical-Conceptual Hierarchy in Speech RepresentationsLinyang He, Qiaolin Wang, Xilin Jiang, Nima MesgaraniEMNLP 2025 · 1 citation
- DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech RepresentationsChao-Hong Tan, Qian Chen, Wen Wang, Chong Deng et al.ICLR 2026 · 8 citations
