MauBERT: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units Discovery
Angelo Ortiz Tandazo, Manel Khentout, Youssef Benchekroun, Thomas Hueber, Emmanuel Dupoux
Abstract
This paper introduces MAUBERT, a multilingual extension of HuBERT that leverages articulatory features for robust cross-lingual phonetic representation learning. We continue Hu-BERT pre-training with supervision based on a phonetic-to-articulatory feature mapping in 55 languages. Our models learn from multilingual data to predict articulatory features or phones, resulting in language-independent representations that capture multilingual phonetic properties. Through comprehensive ABX discriminability testing, we show MAUBERT models produce more context-invariant representations than state-of-the-art multilingual selfsupervised learning models. Additionally, the models effectively adapt to unseen languages and casual speech with minimal self-supervised fine-tuning (10 hours of speech). This establishes an effective approach for instilling linguistic inductive biases in self-supervised speech models. Results. The DiscoPhon results largely corroborate and extend the ABX findings from §5. In zero-shot mode, both MAUBERT variants substantially outperform all baselines across PER, R-value,
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 017dfd38-5efe-4879-80cb-24db74cb56d1Cited by top-tier papers1
Ask how each one uses itBuilds on5
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled DataChengyi Wang, Yu Wu, Yao Qian, Ken'ichi Kumatani et al.ICML 2021 · 140 citations
- UWSpeech: Speech to Speech Translation for Unwritten LanguagesChen Zhang, Xu Tan, Yi Ren, Tao Qin et al.AAAI 2021 · 69 citations
- DinoSR: Self-Distillation and Online Clustering for Self-supervised Speech Representation LearningAlexander H. Liu, Heng-Jui Chang, Michael Auli, Wei-Ning Hsu et al.NeurIPS 2023 · 51 citations
- Towards Robust Speech Representation Learning for Thousands of LanguagesWilliam Chen, Wangyou Zhang, Yifan Peng, Xinjian Li et al.EMNLP 2024 · 19 citations
Related papers
- Identifying Elements Essential for BERT's MultilingualityPhilipp Dufter, Hinrich SchützeEMNLP 2020 · 44 citations
- XLAVS-R: Cross-Lingual Audio-Visual Speech Representation Learning for Noise-Robust Speech PerceptionHyoJung Han, Mohamed Anwar, Juan Pino, Wei-Ning Hsu et al.ACL 2024 · 9 citations
- Cross-Linguistic Syntactic Difference in Multilingual BERT: How Good is It and How Does It Affect Transfer?Ningyu Xu, Tao Gui, Ruotian Ma, Qi Zhang et al.EMNLP 2022 · 4 citations
- Zero-Shot Cross-Lingual Transfer of Neural Machine Translation with Multilingual Pretrained EncodersGuanhua Chen, Shuming Ma, Yun Chen, Li Dong et al.EMNLP 2021 · 30 citations
- Do self-supervised speech models develop human-like perception biases?Juliette Millet, Ewan DunbarACL 2022 · 27 citations
