DinoSR: Self-Distillation and Online Clustering for Self-supervised Speech Representation Learning
Alexander H. Liu, Heng-Jui Chang, Michael Auli, Wei-Ning Hsu, James R. Glass
摘要
In this paper, we introduce self-distillation and online clustering for self-supervised speech representation learning (DinoSR) which combines masked language modeling, self-distillation, and online clustering. We show that these concepts complement each other and result in a strong representation learning model for speech. DinoSR first extracts contextualized embeddings from the input audio with a teacher network, then runs an online clustering system on the embeddings to yield a machine-discovered phone inventory, and finally uses the discretized tokens to guide a student network. We show that DinoSR surpasses previous state-of-the-art performance in several downstream tasks, and provide a detailed analysis of the model and the learned discrete units. Code available at https://github.com/Alexander-H-Liu/dinosr .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- SSDM: Scalable Speech Dysfluency ModelingJiachen Lian, Xuanru Zhou, Zoe Ezzes, Jet Vonk 等NeurIPS 2024 · 被引用 26 次
- Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence LearningApoorv Vyas, Heng-Jui Chang, Cheng-Fu Yang, Po-Yao Huang 等CVPR 2026 · 被引用 23 次
- Towards Robust Speech Representation Learning for Thousands of LanguagesWilliam Chen, Wangyou Zhang, Yifan Peng, Xinjian Li 等EMNLP 2024 · 被引用 19 次
- ES3: Evolving Self-Supervised Learning of Robust Audio-Visual Speech RepresentationsYuanhang Zhang, Shuang Yang, Shiguang Shan, Xilin ChenCVPR 2024 · 被引用 7 次
- StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMsYuhan Song, Linhao Zhang, Chuhan Wu, Aiwei Liu 等ICLR 2026 · 被引用 5 次
它引用的顶会 Paper8
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and LanguageAlexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu 等ICML 2022 · 被引用 1,123 次
- vq-wav2vec: Self-Supervised Learning of Discrete Speech RepresentationsAlexei Baevski, Steffen Schneider, Michael AuliICLR 2020 · 被引用 730 次
相关 Paper
- Self-supervised Semantic-driven Phoneme Discovery for Zero-resource Speech RecognitionLiming Wang, Siyuan Feng, Mark Hasegawa-Johnson, Chang Dong YooACL 2022 · 被引用 5 次
- SyllableLM: Learning Coarse Semantic Units for Speech Language ModelsAlan Baade, Puyuan Peng, David HarwathICLR 2025
- Sylber: Syllabic Embedding Representation of Speech from Raw AudioCheol Jun Cho, Nicholas Lee, Akshat Gupta, Dhruv Agarwal 等ICLR 2025
- Variable-rate hierarchical CPC leads to acoustic unit discovery in speechSantiago Cuervo, Adrian Lancucki, Ricard Marxer, Pawel Rychlikowski 等NeurIPS 2022 · 被引用 20 次
- Self-supervised Neural Factor Analysis for Disentangling Utterance-level Speech RepresentationsWeiwei Lin, Chenhang He, Man-Wai Mak, Youzhi TuICML 2023 · 被引用 6 次
