Speech Vecalign: an Embedding-based Method for Aligning Parallel Speech Documents
Chutong Meng, Philipp Koehn
Abstract
We present Speech Vecalign, a parallel speech document alignment method that monotonically aligns speech segment embeddings and does not depend on text transcriptions. Compared to the baseline method Global Mining (Duquenne et al., 2023a), a variant of speech mining, Speech Vecalign produces longer speech-to-speech alignments. It also demonstrates greater robustness than Local Mining, another speech mining variant, as it produces less noise. We applied Speech Vecalign to 3,000 hours of unlabeled parallel English-German (En-De) speech documents from VoxPopuli, yielding about 1,000 hours of high-quality alignments. We then trained En-De speech-to-speech translation models on the aligned data. Speech Vecalign improves the En-to-De and De-to-En performance over Global Mining by 0.37 and 0.18 ASR-BLEU, respectively. Moreover, our models match or outperform SpeechMatrix model performance, despite using 8 times fewer raw speech documents. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Direct Speech-to-Speech Translation With Discrete UnitsAnn Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu et al.ACL 2022 · 235 citations
- Multimodal and Multilingual Embeddings for Large-Scale Speech MiningPaul-Ambroise Duquenne, Hongyu Gong, Holger SchwenkNeurIPS 2021 · 43 citations
- UnitY: Two-pass Direct Speech-to-speech Translation with Discrete UnitsHirofumi Inaguma, Sravya Popuri, Ilia Kulikov, Peng-Jen Chen et al.ACL 2023 · 30 citations
- SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech TranslationsPaul-Ambroise Duquenne, Hongyu Gong, Ning Dong, Jingfei Du et al.ACL 2023 · 18 citations
Related papers
- VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and InterpretationChanghan Wang, Morgane Rivière, Ann Lee, Anne Wu et al.ACL 2021
- Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference OptimizationChaoqun Cui, Liangbin Huang, Shijing Wang, Zhe Tong et al.ACL 2025
- Discrete Cross-Modal Alignment Enables Zero-Shot Speech TranslationChen Wang, Yuchen Liu, Boxing Chen, Jiajun Zhang et al.EMNLP 2022 · 3 citations
- Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech SynthesisWeiwei Lin, Chenhang HeICLR 2025
- Simple and Effective Unsupervised Speech TranslationChanghan Wang, Hirofumi Inaguma, Peng-Jen Chen, Ilia Kulikov et al.ACL 2023 · 9 citations
