Do self-supervised speech models develop human-like perception biases?
Juliette Millet, Ewan Dunbar
Abstract
Self-supervised models for speech processing form representational spaces without using any external labels. Increasingly, they appear to be a feasible way of at least partially eliminating costly manual annotations, a problem of particular concern for low-resource languages. But what kind of representational spaces do these models construct?Human perception specializes to the sounds of listeners’ native languages. Does the same thing happen in self-supervised models? We examine the representational spaces of three kinds of state of the art self-supervised models: wav2vec, HuBERT and contrastive predictive coding (CPC), and compare them with the perceptual spaces of French-speaking and English-speaking human listeners, both globally and taking account of the behavioural differences between the two language groups. We show that the CPC model shows a small native language effect, but that wav2vec and HuBERT seem to develop a universal speech perception space which is not language specific. A comparison against the predictions of supervised phone recognisers suggests that all three self-supervised models capture relatively fine-grained perceptual phenomena, while supervised models are better at capturing coarser, phone-level effects, and effects of listeners’ native language, on perception.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4bae4117-db76-498a-a50d-19ea0d1ca301Cited by top-tier papers2
- Toward a realistic model of speech processing in the brain with self-supervised learningJuliette Millet, Charlotte Caucheteux, Pierre Orhan, Yves Boubenec et al.NeurIPS 2022 · 164 citations
- Homophone Disambiguation Reveals Patterns of Context Mixing in Speech TransformersHosein Mohebbi, Grzegorz Chrupala, Willem H. Zuidema, Afra AlishahiEMNLP 2023 · 1 citation
Builds on2
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and InterpretationChanghan Wang, Morgane Rivière, Ann Lee, Anne Wu et al.ACL 2021
Related papers
- Self-Supervised Models of Audio Effectively Explain Human Cortical Responses to SpeechAditya R. Vaidya, Shailee Jain, Alexander HuthICML 2022 · 81 citations
- Self-supervised Neural Factor Analysis for Disentangling Utterance-level Speech RepresentationsWeiwei Lin, Chenhang He, Man-Wai Mak, Youzhi TuICML 2023 · 6 citations
- From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation modelMarvin Lavechin, Thomas HueberEMNLP 2025
- Unsupervised Speech RecognitionAlexei Baevski, Wei-Ning Hsu, Alexis Conneau, Michael AuliNeurIPS 2021 · 309 citations
- RepCodec: A Speech Representation Codec for Speech TokenizationZhichao Huang, Chutong Meng, Tom KoACL 2024 · 16 citations
