SpidR-Adapt: A Universal Speech Representation Model for Few-Shot Adaptation
Mahi Luthra, Jiayi Shen, Maxime Poli, Angelo Ortiz Tandazo, Yosuke Higuchi, Youssef Benchekroun, Martin Gleize, Charles-Éric Saint-James, Dongyan Lin, Phillip Rust, Angel Villar-Corrales, Surya Parimi
Abstract
Human infants, with only a few hundred hours of speech exposure, acquire basic units of new languages, highlighting a striking efficiency gap compared to the data-hungry self-supervised speech models. To address this gap, this paper introduces SpidR-Adapt for rapid adaptation of speech units to new languages using minimal unlabeled data. We cast such low-resource speech representation learning as a meta-learning problem and construct a multi-task adaptive pre-training (MAdaPT) protocol which formulates the adaptation process as a bi-level optimization framework. To enable scalable meta-training under this framework, we propose a novel heuristic solution, first-order bi-level optimization (FOBLO), avoiding heavy computation costs. Finally, we stabilize meta-training by using a robust initialization through interleaved supervision which alternates self-supervised and supervised objectives. Empirically, SpidR-Adapt achieves rapid gains in phonemic discriminability (ABX) and downstream spoken language modeling scores (sWUGGY, sBLIMP, tSC), surpassing in-domain toplines after training on less than 1h of target-language audio and delivering greater data efficiency than standard multi-task training. These findings highlight a practical, architecture-agnostic path toward biologically inspired, data-efficient representations. We open-source the training code and model checkpoints at https://github.com/facebookresearch/spidr-adapt.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fcc24024-0e62-4a89-9706-e3ee94338125Builds on7
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- DinoSR: Self-Distillation and Online Clustering for Self-supervised Speech Representation LearningAlexander H. Liu, Heng-Jui Chang, Michael Auli, Wei-Ning Hsu et al.NeurIPS 2023 · 51 citations
- Improving Language Plasticity via Pretraining with Active ForgettingYihong Chen, Kelly Marchisio, Roberta Raileanu, David Ifeoluwa Adelani et al.NeurIPS 2023 · 49 citations
- MauBERT: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units DiscoveryAngelo Ortiz Tandazo, Manel Khentout, Youssef Benchekroun, Thomas Hueber et al.ACL 2026 · 1 citation
- Fast and Robust: Task Sampling with Posterior and Diversity Synergies for Adaptive Decision-Makers in Randomized EnvironmentsYun Qu, Cheems Wang, Yixiu Mao, Yiqin Lv et al.ICML 2025
Related papers
- MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual TransferJonas Pfeiffer, Ivan Vulic, Iryna Gurevych, Sebastian RuderEMNLP 2020 · 36 citations
- Self-Supervised Meta-Learning for Few-Shot Natural Language Classification TasksTrapit Bansal, Rishikesh Jha, Tsendsuren Munkhdalai, Andrew McCallumEMNLP 2020 · 9 citations
- Learn to Cross-lingual Transfer with Meta Graph Learning Across Heterogeneous LanguagesZheng Li, Mukul Kumar, William Headden, Bing Yin et al.EMNLP 2020 · 26 citations
- Rapid Word Learning Through Meta In-Context LearningWentao Wang, Guangyuan Jiang, Tal Linzen, Brenden M. LakeEMNLP 2025
- Adversarial Meta Sampling for Multilingual Low-Resource Speech RecognitionYubei Xiao, Ke Gong, Pan Zhou, Guolin Zheng et al.AAAI 2021 · 37 citations
