Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction
Jiatong Shi, Hirofumi Inaguma, Xutai Ma, Ilia Kulikov, Anna Y. Sun
摘要
Existing Self-Supervised Learning (SSL) models for speech typically process speech signals at a fixed resolution of 20 milliseconds. This approach overlooks the varying informational content present at different resolutions in speech signals. In contrast, this paper aims to incorporate multi-resolution information into speech self-supervised representation learning. We introduce a SSL model that leverages a hierarchical Transformer architecture, complemented by HuBERT-style masked prediction objectives, to process speech at multiple resolutions. Experimental results indicate that the proposed model not only achieves more efficient inference but also exhibits superior or comparable performance to the original Hu-BERT model over various tasks. Specifically, significant performance improvements over the original HuBERT have been observed in fine-tuning experiments on the LibriSpeech speech recognition benchmark as well as in evaluations using the Speech Universal PERformance Benchmark (SUPERB) and Multilingual SUPERB (ML-SUPERB). * The work was conducted by Jiatong Shi during his summer internship at Meta.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- SSDM: Scalable Speech Dysfluency ModelingJiachen Lian, Xuanru Zhou, Zoe Ezzes, Jet Vonk 等NeurIPS 2024 · 被引用 26 次
- Towards Robust Speech Representation Learning for Thousands of LanguagesWilliam Chen, Wangyou Zhang, Yifan Peng, Xinjian Li 等EMNLP 2024 · 被引用 19 次
- SPEAR: A Unified SSL Framework for Learning Speech and Audio RepresentationsXiaoyu Yang, Yifan Yang, Zengrui Jin, Ziyun Cui 等ICML 2026 · 被引用 12 次
- DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language ModelsYuanyuan Wang, Dongchao Yang, Yiwen Shao, Hangting Chen 等AAAI 2026 · 被引用 3 次
- An Exploration of Mamba for Speech Self-Supervised ModelsTzu-Quan Lin, Heng-Cheng Kuo, Tzu-Chieh Wei, Hsi-Chun Cheng 等ACL 2026 · 被引用 2 次
它引用的顶会 Paper14
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and LanguageAlexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu 等ICML 2022 · 被引用 1,123 次
- MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised TrainingYizhi Li, Ruibin Yuan, Ge Zhang, Yinghao Ma 等ICLR 2024 · 被引用 277 次
- Self-supervised learning with random-projection quantizer for speech recognitionChung-Cheng Chiu, James Qin, Yu Zhang, Jiahui Yu 等ICML 2022 · 被引用 245 次
相关 Paper
- Self-supervised Neural Factor Analysis for Disentangling Utterance-level Speech RepresentationsWeiwei Lin, Chenhang He, Man-Wai Mak, Youzhi TuICML 2023 · 被引用 6 次
- Self-Supervised Models of Audio Effectively Explain Human Cortical Responses to SpeechAditya R. Vaidya, Shailee Jain, Alexander HuthICML 2022 · 被引用 81 次
- Sylber: Syllabic Embedding Representation of Speech from Raw AudioCheol Jun Cho, Nicholas Lee, Akshat Gupta, Dhruv Agarwal 等ICLR 2025
- ContentVec: An Improved Self-Supervised Speech Representation by Disentangling SpeakersKaizhi Qian, Yang Zhang, Heting Gao, Junrui Ni 等ICML 2022 · 被引用 157 次
- CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech ProcessingYen-Ju Lu, Jing Liu, Thomas Thebaud, Laureano Moro-Velázquez 等NeurIPS 2024 · 被引用 5 次
