SpikeVoice: High-Quality Text-to-Speech Via Efficient Spiking Neural Network
Kexin Wang, Jiahong Zhang, Yong Ren, Man Yao, Di Shang, Bo Xu, Guoqi Li
Abstract
Brain-inspired Spiking Neural Network (SNN) has demonstrated its effectiveness and efficiency in vision, natural language, and speech understanding tasks, indicating their capacity to "see", "listen", and "read". In this paper, we design SpikeVoice, which performs high-quality Text-To-Speech (TTS) via SNN, to explore the potential of SNN to "speak". A major obstacle to using SNN for such generative tasks lies in the demand for models to grasp long-term dependencies. The serial nature of spiking neurons, however, leads to the invisibility of information at future spiking time steps, limiting SNN models to capture sequence dependencies solely within the same time step. We term this phenomenon "partial-time dependency". To address this issue, we introduce Spiking Temporal-Sequential Attention (STSA) in the SpikeVoice. To the best of our knowledge, SpikeVoice is the first TTS work in the SNN field. We perform experiments using four wellestablished datasets that cover both Chinese and English languages, encompassing scenarios with both single-speaker and multi-speaker configurations. The results demonstrate that SpikeVoice can achieve results comparable to Artificial Neural Networks (ANN) with only 10.5% energy consumption of ANN. Both our demo and code are available as supplementary material.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 45cdb53c-1984-4329-b41c-0f2992ce4f93Cited by top-tier papers2
- MMDEND: Dendrite-Inspired Multi-Branch Multi-Compartment Parallel Spiking Neuron for Sequence ModelingKexin Wang, Yuhong Chou, Di Shang, Shijie Mei et al.ACL 2025 · 4 citations
- Towards Lossless Memory-efficient Training of Spiking Neural Networks via Gradient Checkpointing and Spike CompressionYifan Huang, Wei Fang, Zecheng Hao, Zhengyu Ma et al.ICLR 2026
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
Related papers
- Efficient and Effective Time-Series Forecasting with Spiking Neural NetworksChangze Lv, Yansen Wang, Dongqi Han, Xiaoqing Zheng et al.ICML 2024 · 29 citations
- SpikeZIP-TF: Conversion is All You Need for Transformer-based SNNKang You, Zekai Xu, Chen Nie, Zhijie Deng et al.ICML 2024 · 20 citations
- SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command RecognitionJiaqi Wang, Liutao Yu, Xiongri Shen, Sihang Guo et al.AAAI 2026 · 1 citation
- Temporal Dynamics Enhancer for Directly Trained Spiking Object DetectorsFan Luo, Zeyu Gao, Xinhao Luo, Kai Zhao et al.AAAI 2026
- SpikF: Spiking Fourier Network for Efficient Long-term PredictionWenjie Wu, Dexuan Huo, Hong ChenICML 2025
