SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech Translations
Paul-Ambroise Duquenne, Hongyu Gong, Ning Dong, Jingfei Du, Ann Lee, Vedanuj Goswami, Changhan Wang, Juan Pino, Benoît Sagot, Holger Schwenk
Abstract
We present SpeechMatrix, a large-scale multilingual corpus of speech-to-speech translations (S2ST) mined from real speech of European Parliament recordings. It contains speech alignments in 136 language pairs with a total of 418 thousand hours of speech. To evaluate the quality of this parallel speech, we train bilingual speech-to-speech translation models on mined data only and establish extensive baseline results on Europarl-ST, VoxPopuli and FLEURS test sets. Enabled by the multilinguality of SpeechMatrix, we also explore multilingual speech-to-speech translation, a topic which was addressed by few other works. We also demonstrate that model pre-training and sparse scaling using Mixture-of-Experts bring large gains to translation performance. We are open-sourcing the mined data, speech encoders used for mining, multilingual HuBERT models in four language families for target unit generation, language-specific vocoders for speech synthesis from discrete units, and S2S models trained and presented in this work. 1 * Equal contributions 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- TransVIP: Speech to Speech Translation System with Voice and Isochrony PreservationChenyang Le, Yao Qian, Dongmei Wang, Long Zhou et al.NeurIPS 2024 · 25 citations
- AV-TranSpeech: Audio-Visual Robust Speech-to-Speech TranslationRongjie Huang, Huadai Liu, Xize Cheng, Yi Ren et al.ACL 2023 · 9 citations
- BLASER: A Text-Free Speech-to-Speech Translation Evaluation MetricMingda Chen, Paul-Ambroise Duquenne, Pierre Andrews, Justine Kao et al.ACL 2023 · 8 citations
- Make-A-Voice: Revisiting Voice Large Language Models as Scalable Multilingual and Multitask LearnersRongjie Huang, Chunlei Zhang, Yongqi Wang, Dongchao Yang et al.ACL 2024 · 6 citations
- Towards Building Large Scale Datasets and State-of-the-Art Automatic Speech Translation Systems for 14 Indian LanguagesAshwin Sankar, Sparsh Jain, Nikhil Narasimhan, Devilal Choudhary et al.ACL 2025 · 5 citations
Builds on8
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- BASE Layers: Simplifying Training of Large, Sparse ModelsMike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal et al.ICML 2021 · 382 citations
- Direct Speech-to-Speech Translation With Discrete UnitsAnn Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu et al.ACL 2022 · 235 citations
- Making Monolingual Sentence Embeddings Multilingual using Knowledge DistillationNils Reimers, Iryna GurevychEMNLP 2020 · 54 citations
Related papers
- Multimodal and Multilingual Embeddings for Large-Scale Speech MiningPaul-Ambroise Duquenne, Hongyu Gong, Holger SchwenkNeurIPS 2021 · 43 citations
- VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and InterpretationChanghan Wang, Morgane Rivière, Ann Lee, Anne Wu et al.ACL 2021
- Speech Vecalign: an Embedding-based Method for Aligning Parallel Speech DocumentsChutong Meng, Philipp KoehnEMNLP 2025
- From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech RecognitionTianduo Wang, Lu Xu, Wei Lu, Shanbo ChengEMNLP 2025 · 1 citation
- An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)Laurie Burchell, Ona De Gibert Bonet, Nikolay Arefyev, Mikko Aulamo et al.ACL 2025
