DinoSR: Self-Distillation and Online Clustering for Self-supervised Speech Representation Learning
Alexander H. Liu, Heng-Jui Chang, Michael Auli, Wei-Ning Hsu, James R. Glass
Abstract
In this paper, we introduce self-distillation and online clustering for self-supervised speech representation learning (DinoSR) which combines masked language modeling, self-distillation, and online clustering. We show that these concepts complement each other and result in a strong representation learning model for speech. DinoSR first extracts contextualized embeddings from the input audio with a teacher network, then runs an online clustering system on the embeddings to yield a machine-discovered phone inventory, and finally uses the discretized tokens to guide a student network. We show that DinoSR surpasses previous state-of-the-art performance in several downstream tasks, and provide a detailed analysis of the model and the learned discrete units. Code available at https://github.com/Alexander-H-Liu/dinosr .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1bf0db34-bd19-4fcf-84bb-1fb169fbcd46Cited by top-tier papers12
- SSDM: Scalable Speech Dysfluency ModelingJiachen Lian, Xuanru Zhou, Zoe Ezzes, Jet Vonk et al.NeurIPS 2024 · 26 citations
- Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence LearningApoorv Vyas, Heng-Jui Chang, Cheng-Fu Yang, Po-Yao Huang et al.CVPR 2026 · 23 citations
- Towards Robust Speech Representation Learning for Thousands of LanguagesWilliam Chen, Wangyou Zhang, Yifan Peng, Xinjian Li et al.EMNLP 2024 · 19 citations
- ES3: Evolving Self-Supervised Learning of Robust Audio-Visual Speech RepresentationsYuanhang Zhang, Shuang Yang, Shiguang Shan, Xilin ChenCVPR 2024 · 7 citations
- StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMsYuhan Song, Linhao Zhang, Chuhan Wu, Aiwei Liu et al.ICLR 2026 · 5 citations
Builds on8
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and LanguageAlexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu et al.ICML 2022 · 1,123 citations
- vq-wav2vec: Self-Supervised Learning of Discrete Speech RepresentationsAlexei Baevski, Steffen Schneider, Michael AuliICLR 2020 · 730 citations
Related papers
- Self-supervised Semantic-driven Phoneme Discovery for Zero-resource Speech RecognitionLiming Wang, Siyuan Feng, Mark Hasegawa-Johnson, Chang Dong YooACL 2022 · 5 citations
- SyllableLM: Learning Coarse Semantic Units for Speech Language ModelsAlan Baade, Puyuan Peng, David HarwathICLR 2025
- Sylber: Syllabic Embedding Representation of Speech from Raw AudioCheol Jun Cho, Nicholas Lee, Akshat Gupta, Dhruv Agarwal et al.ICLR 2025
- Variable-rate hierarchical CPC leads to acoustic unit discovery in speechSantiago Cuervo, Adrian Lancucki, Ricard Marxer, Pawel Rychlikowski et al.NeurIPS 2022 · 20 citations
- Self-supervised Neural Factor Analysis for Disentangling Utterance-level Speech RepresentationsWeiwei Lin, Chenhang He, Man-Wai Mak, Youzhi TuICML 2023 · 6 citations
