Speech Vecalign: an Embedding-based Method for Aligning Parallel Speech Documents
Chutong Meng, Philipp Koehn
摘要
We present Speech Vecalign, a parallel speech document alignment method that monotonically aligns speech segment embeddings and does not depend on text transcriptions. Compared to the baseline method Global Mining (Duquenne et al., 2023a), a variant of speech mining, Speech Vecalign produces longer speech-to-speech alignments. It also demonstrates greater robustness than Local Mining, another speech mining variant, as it produces less noise. We applied Speech Vecalign to 3,000 hours of unlabeled parallel English-German (En-De) speech documents from VoxPopuli, yielding about 1,000 hours of high-quality alignments. We then trained En-De speech-to-speech translation models on the aligned data. Speech Vecalign improves the En-to-De and De-to-En performance over Global Mining by 0.37 and 0.18 ASR-BLEU, respectively. Moreover, our models match or outperform SpeechMatrix model performance, despite using 8 times fewer raw speech documents. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Direct Speech-to-Speech Translation With Discrete UnitsAnn Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu 等ACL 2022 · 被引用 235 次
- Multimodal and Multilingual Embeddings for Large-Scale Speech MiningPaul-Ambroise Duquenne, Hongyu Gong, Holger SchwenkNeurIPS 2021 · 被引用 43 次
- UnitY: Two-pass Direct Speech-to-speech Translation with Discrete UnitsHirofumi Inaguma, Sravya Popuri, Ilia Kulikov, Peng-Jen Chen 等ACL 2023 · 被引用 30 次
- SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech TranslationsPaul-Ambroise Duquenne, Hongyu Gong, Ning Dong, Jingfei Du 等ACL 2023 · 被引用 18 次
相关 Paper
- VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and InterpretationChanghan Wang, Morgane Rivière, Ann Lee, Anne Wu 等ACL 2021
- Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference OptimizationChaoqun Cui, Liangbin Huang, Shijing Wang, Zhe Tong 等ACL 2025
- Discrete Cross-Modal Alignment Enables Zero-Shot Speech TranslationChen Wang, Yuchen Liu, Boxing Chen, Jiajun Zhang 等EMNLP 2022 · 被引用 3 次
- Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech SynthesisWeiwei Lin, Chenhang HeICLR 2025
- Simple and Effective Unsupervised Speech TranslationChanghan Wang, Hirofumi Inaguma, Peng-Jen Chen, Ilia Kulikov 等ACL 2023 · 被引用 9 次
