SyllableLM: Learning Coarse Semantic Units for Speech Language Models
Alan Baade, Puyuan Peng, David Harwath
摘要
Language models require tokenized inputs. However, tokenization strategies for continuous data like audio and vision are often based on simple heuristics such as fixed sized convolutions or discrete clustering, which do not necessarily align with the semantic structure of the data. For speech in particular, the high resolution of waveforms (16,000 samples/second or more) presents a significant challenge as speech-based language models have had to use several times more tokens per word than text-based language models. In this work, we introduce a controllable self-supervised technique to merge speech representations into coarser syllable-like units while still preserving semantic information. We do this by 1) extracting noisy boundaries through analyzing correlations in pretrained encoder losses and 2) iteratively improving model representations with a novel distillation technique. Our method produces controllable-rate semantic units at as low as 5Hz and 60bps and achieves SotA in syllabic segmentation and clustering. Using these coarse tokens, we successfully train SyllableLM, a Speech Language Model (SpeechLM) that matches or outperforms current SotA SpeechLMs on a range of spoken language modeling tasks. SyllableLM also achieves significant improvements in efficiency with a 30x reduction in training compute and a 4x wall-clock inference speedup.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI FeedbackGuan-Ting Lin, Prashanth Gurunath Shivakumar, Aditya Gourav, Yile Gu 等ACL 2025 · 被引用 33 次
- FlexiCodec: A Dynamic Neural Audio Codec for Low Frame RatesJiaqi Li, Yao Qian, Yuxuan Hu, leying zhang 等ICLR 2026 · 被引用 27 次
- Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration CodingRui-Chen Zheng, Wenrui Liu, Hui-Peng Du, Qinglin Zhang 等AAAI 2026 · 被引用 4 次
- Speech Token Prediction via Compressed-to-fine Language Modeling for Speech GenerationWenrui Liu, Qian Chen, Wen Wang, Guanrou Yang 等ACM MM 2025
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and LanguageAlexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu 等ICML 2022 · 被引用 1,123 次
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar 等NeurIPS 2023 · 被引用 910 次
相关 Paper
- Sylber: Syllabic Embedding Representation of Speech from Raw AudioCheol Jun Cho, Nicholas Lee, Akshat Gupta, Dhruv Agarwal 等ICLR 2025
- Scaling Properties of Speech Language ModelsSantiago Cuervo, Ricard MarxerEMNLP 2024 · 被引用 5 次
- Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMsDingdong Wang, Junan Li, Mingyu Cui, Dongchao Yang 等EMNLP 2025 · 被引用 1 次
- DinoSR: Self-Distillation and Online Clustering for Self-supervised Speech Representation LearningAlexander H. Liu, Heng-Jui Chang, Michael Auli, Wei-Ning Hsu 等NeurIPS 2023 · 被引用 51 次
- Generative Spoken Language Model based on continuous word-sized audio tokensRobin Algayres, Yossi Adi, Tu Anh Nguyen, Jade Copet 等EMNLP 2023 · 被引用 3 次
