Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer
Yongxin Zhu, Dan Su, Liqiang He, Linli Xu, Dong Yu
Abstract
While recent advancements in speech language models have achieved significant progress, they face remarkable challenges in modeling the long acoustic sequences of neural audio codecs. In this paper, we introduce Generative Pretrained Speech Transformer (GPST), a hierarchical transformer designed for efficient speech language modeling. GPST quantizes audio waveforms into two distinct types of discrete speech representations and integrates them within a hierarchical transformer architecture, allowing for a unified one-stage generation process and enhancing Hi-Res audio generation capabilities. By training on large corpora of speeches in an end-to-end unsupervised manner, GPST can generate syntactically consistent speech with diverse speaker identities. Given a brief 3-second prompt, GPST can produce natural and coherent personalized speech, demonstrating in-context learning abilities. Moreover, our approach can be easily extended to spoken cross-lingual speech generation by incorporating multi-lingual semantic tokens and universal acoustic tokens. Experimental results indicate that GPST significantly outperforms the existing speech language models in terms of word error rate, speech quality, and speaker similarity. See https://youngsheen.github. io/GPST/demo for demo samples.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ebc52d4a-f0cb-4e7d-82a0-17b855730255Cited by top-tier papers4
- MIBURI: Towards Expressive Interactive Gesture SynthesisMuhammad Hamza Mughal, Rishabh Dabral, Vera Demberg, Christian TheobaltCVPR 2026 · 10 citations
- Addressing Representation Collapse in Vector Quantized Models with One Linear LayerYongxin Zhu, Bocheng Li, Yifei Xin, Zhihua Xia et al.ICCV 2025 · 5 citations
- Efficient Speech Language Modeling via Energy Distance in Continuous Latent SpaceZhengrui Ma, Yang Feng, Chenze Shao, Fandong Meng et al.NeurIPS 2025 · 5 citations
- Recent Advances in Speech Language Models: A SurveyWenqian Cui, Dianzhi Yu, Xiaoqi Jiao, Ziqiao Meng et al.ACL 2025
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for EveryoneEdresson Casanova, Julian Weber, Christopher Dane Shulby, Arnaldo Cândido Júnior et al.ICML 2022 · 602 citations
- Direct Speech-to-Speech Translation With Discrete UnitsAnn Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu et al.ACL 2022 · 235 citations
- MEGABYTE: Predicting Million-byte Sequences with Multiscale TransformersLili Yu, Daniel Simig, Colin Flaherty, Armen Aghajanyan et al.NeurIPS 2023 · 197 citations
Related papers
- Generative Pretrained Structured Transformers: Unsupervised Syntactic Language Models at ScaleXiang Hu, Pengyu Ji, Qingyang Zhu, Wei Wu et al.ACL 2024 · 1 citation
- GenSE: Generative Speech Enhancement via Language Models using Hierarchical ModelingJixun Yao, Hexin Liu, Chen Chen, Yuchen Hu et al.ICLR 2025
- Hierarchical Semantic-Acoustic Modeling via Semi-Discrete Residual Representations for Expressive End-to-End Speech SynthesisYixuan Zhou, Guoyang Zeng, Xin Liu, Xiang Li et al.ICLR 2026
- Scaling Transformers for End-to-End Discrete Audio TokenizationYitian Gong, Kuangwei Chen, Zhaoye Fei, Xiaogui Yang et al.ICML 2026
- Speech Token Prediction via Compressed-to-fine Language Modeling for Speech GenerationWenrui Liu, Qian Chen, Wen Wang, Guanrou Yang et al.ACM MM 2025
