Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction
Jiatong Shi, Hirofumi Inaguma, Xutai Ma, Ilia Kulikov, Anna Y. Sun
Abstract
Existing Self-Supervised Learning (SSL) models for speech typically process speech signals at a fixed resolution of 20 milliseconds. This approach overlooks the varying informational content present at different resolutions in speech signals. In contrast, this paper aims to incorporate multi-resolution information into speech self-supervised representation learning. We introduce a SSL model that leverages a hierarchical Transformer architecture, complemented by HuBERT-style masked prediction objectives, to process speech at multiple resolutions. Experimental results indicate that the proposed model not only achieves more efficient inference but also exhibits superior or comparable performance to the original Hu-BERT model over various tasks. Specifically, significant performance improvements over the original HuBERT have been observed in fine-tuning experiments on the LibriSpeech speech recognition benchmark as well as in evaluations using the Speech Universal PERformance Benchmark (SUPERB) and Multilingual SUPERB (ML-SUPERB). * The work was conducted by Jiatong Shi during his summer internship at Meta.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d8541a4d-3081-4f2e-bde7-c9153a828f3fCited by top-tier papers5
- SSDM: Scalable Speech Dysfluency ModelingJiachen Lian, Xuanru Zhou, Zoe Ezzes, Jet Vonk et al.NeurIPS 2024 · 26 citations
- Towards Robust Speech Representation Learning for Thousands of LanguagesWilliam Chen, Wangyou Zhang, Yifan Peng, Xinjian Li et al.EMNLP 2024 · 19 citations
- SPEAR: A Unified SSL Framework for Learning Speech and Audio RepresentationsXiaoyu Yang, Yifan Yang, Zengrui Jin, Ziyun Cui et al.ICML 2026 · 12 citations
- DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language ModelsYuanyuan Wang, Dongchao Yang, Yiwen Shao, Hangting Chen et al.AAAI 2026 · 3 citations
- An Exploration of Mamba for Speech Self-Supervised ModelsTzu-Quan Lin, Heng-Cheng Kuo, Tzu-Chieh Wei, Hsi-Chun Cheng et al.ACL 2026 · 2 citations
Builds on14
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and LanguageAlexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu et al.ICML 2022 · 1,123 citations
- MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised TrainingYizhi Li, Ruibin Yuan, Ge Zhang, Yinghao Ma et al.ICLR 2024 · 277 citations
- Self-supervised learning with random-projection quantizer for speech recognitionChung-Cheng Chiu, James Qin, Yu Zhang, Jiahui Yu et al.ICML 2022 · 245 citations
Related papers
- Self-supervised Neural Factor Analysis for Disentangling Utterance-level Speech RepresentationsWeiwei Lin, Chenhang He, Man-Wai Mak, Youzhi TuICML 2023 · 6 citations
- Self-Supervised Models of Audio Effectively Explain Human Cortical Responses to SpeechAditya R. Vaidya, Shailee Jain, Alexander HuthICML 2022 · 81 citations
- Sylber: Syllabic Embedding Representation of Speech from Raw AudioCheol Jun Cho, Nicholas Lee, Akshat Gupta, Dhruv Agarwal et al.ICLR 2025
- ContentVec: An Improved Self-Supervised Speech Representation by Disentangling SpeakersKaizhi Qian, Yang Zhang, Heting Gao, Junrui Ni et al.ICML 2022 · 157 citations
- CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech ProcessingYen-Ju Lu, Jing Liu, Thomas Thebaud, Laureano Moro-Velázquez et al.NeurIPS 2024 · 5 citations
