Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
Dingdong Wang, Junan Li, Mingyu Cui, Dongchao Yang, Xueyuan Chen, Helen M. Meng
摘要
With the rise of Speech Large Language Models (SpeechLLMs), two dominant approaches have emerged for speech processing: discrete tokens and continuous features. Each approach has demonstrated strong capabilities in audiorelated processing tasks. However, the performance gap between these two paradigms has not been thoroughly explored. To address this gap, we present a fair comparison of selfsupervised learning (SSL)-based discrete and continuous features under the same experimental settings. We evaluate their performance across six spoken language understandingrelated tasks using both small and large-scale LLMs (Qwen1.5-0.5B and Llama3.1-8B). We further conduct in-depth analyses, including efficient comparison, SSL layer analysis, LLM layer analysis, and robustness comparison. Our findings reveal that continuous features generally outperform discrete tokens in various tasks. Each speech processing method exhibits distinct characteristics and patterns in how it learns and processes speech information. We hope our findings will provide valuable insights to advance spoken language understanding in SpeechLLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- SALMONN: Towards Generic Hearing Abilities for Large Language ModelsChangli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen 等ICLR 2024 · 被引用 557 次
- InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-trainingDingdong Wang, Jin Xu, Ruihang Chu, Zhifang Guo 等ACL 2025 · 被引用 9 次
- Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language ModelsSander Land, Max BartoloEMNLP 2024 · 被引用 4 次
- ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language ModelingDongchao Yang, Songxiang Liu, Haohan Guo, Jiankun Zhao 等ICML 2025
相关 Paper
- A Variational Framework for Improving Naturalness in Generative Spoken Language ModelsLi-Wei Chen, Takuya Higuchi, Zakaria Aldeneh, Ahmed Hussen Abdelaziz 等ICML 2025
- Generative Spoken Language Model based on continuous word-sized audio tokensRobin Algayres, Yossi Adi, Tu Anh Nguyen, Jade Copet 等EMNLP 2023 · 被引用 3 次
- SyllableLM: Learning Coarse Semantic Units for Speech Language ModelsAlan Baade, Puyuan Peng, David HarwathICLR 2025
- Layer-wise Minimal Pair Probing Reveals Contextual Grammatical-Conceptual Hierarchy in Speech RepresentationsLinyang He, Qiaolin Wang, Xilin Jiang, Nima MesgaraniEMNLP 2025 · 被引用 1 次
- DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech RepresentationsChao-Hong Tan, Qian Chen, Wen Wang, Chong Deng 等ICLR 2026 · 被引用 8 次
