Self-supervised Neural Factor Analysis for Disentangling Utterance-level Speech Representations
Weiwei Lin, Chenhang He, Man-Wai Mak, Youzhi Tu
Abstract
Self-supervised learning (SSL) speech models such as wav2vec and HuBERT have demonstrated state-of-the-art performance on automatic speech recognition (ASR) and proved to be extremely useful in low label-resource settings. However, the success of SSL models has yet to transfer to utterance-level tasks such as speaker, emotion, and language recognition, which still require supervised fine-tuning of the SSL models to obtain good performance. We argue that the problem is caused by the lack of disentangled representations and an utterance-level learning objective for these tasks. Inspired by how HuBERT uses clustering to discover hidden acoustic units, we formulate a factor analysis (FA) model that uses the discovered hidden acoustic units to align the SSL features. The underlying utterance-level representations are disentangled from the content of speech using probabilistic inference on the aligned features. Furthermore, the variational lower bound derived from the FA model provides an utterance-level objective, allowing error gradients to be backpropagated to the Transformer layers to learn highly discriminative acoustic units. When used in conjunction with HuBERT's masked prediction training, our models outperform the current best model, WavLM, on all utterance-level non-semantic tasks on the SUPERB benchmark with only 20% of labeled data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2e0c1781-d69e-4436-aae3-c55729b6a329Builds on5
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- vq-wav2vec: Self-Supervised Learning of Discrete Speech RepresentationsAlexei Baevski, Steffen Schneider, Michael AuliICLR 2020 · 730 citations
- ContentVec: An Improved Self-Supervised Speech Representation by Disentangling SpeakersKaizhi Qian, Yang Zhang, Heting Gao, Junrui Ni et al.ICML 2022 · 157 citations
Related papers
- Self-Supervised Models of Audio Effectively Explain Human Cortical Responses to SpeechAditya R. Vaidya, Shailee Jain, Alexander HuthICML 2022 · 81 citations
- Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit PredictionJiatong Shi, Hirofumi Inaguma, Xutai Ma, Ilia Kulikov et al.ICLR 2024 · 39 citations
- SPEAR: A Unified SSL Framework for Learning Speech and Audio RepresentationsXiaoyu Yang, Yifan Yang, Zengrui Jin, Ziyun Cui et al.ICML 2026 · 12 citations
- Speech Self-Supervised Learning Using Diffusion Model Synthetic DataHeting Gao, Kaizhi Qian, Junrui Ni, Chuang Gan et al.ICML 2024 · 8 citations
- An Exploration of Mamba for Speech Self-Supervised ModelsTzu-Quan Lin, Heng-Cheng Kuo, Tzu-Chieh Wei, Hsi-Chun Cheng et al.ACL 2026 · 2 citations
