Multimodal and Multilingual Embeddings for Large-Scale Speech Mining
Paul-Ambroise Duquenne, Hongyu Gong, Holger Schwenk
摘要
We present an approach to encode a speech signal into a fixed-size representation which minimizes the cosine loss with the existing massively multilingual LASER text embedding space. Sentences are close in this embedding space, independently of their language and modality, either text or audio. Using a similarity metric in that multimodal embedding space, we perform mining of audio in German, French, Spanish and English from Librivox against billions of sentences from Common Crawl. This yielded more than twenty thousand hours of aligned speech translations. To evaluate the automatically mined speech/text corpora, we train neural speech translation systems for several languages pairs. Adding the mined data, achieves significant improvements in the BLEU score on the CoVoST2 and the MUST-C test sets with respect to a very competitive baseline. Our approach can also be used to directly perform speech-to-speech mining, without the need to first transcribe or translate the data. We obtain more than one thousand three hundred hours of aligned speech in French, German, Spanish and English. This speech corpus has the potential to boost research in speech-to-speech translation which suffers from scarcity of natural end-to-end training data. All the mined multimodal corpora will be made freely available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech TranslationsPaul-Ambroise Duquenne, Hongyu Gong, Ning Dong, Jingfei Du 等ACL 2023 · 被引用 18 次
- RegBN: Batch Normalization of Multimodal Data with RegularizationMorteza Ghahremani, Christian WachingerNeurIPS 2023 · 被引用 15 次
- BLASER: A Text-Free Speech-to-Speech Translation Evaluation MetricMingda Chen, Paul-Ambroise Duquenne, Pierre Andrews, Justine Kao 等ACL 2023 · 被引用 8 次
- T-Modules: Translation Modules for Zero-Shot Cross-Modal Machine TranslationPaul-Ambroise Duquenne, Hongyu Gong, Benoît Sagot, Holger SchwenkEMNLP 2022 · 被引用 8 次
- DiffS2UT: A Semantic Preserving Diffusion Model for Textless Direct Speech-to-Speech TranslationYongxin Zhu, Zhujin Gao, Xinyuan Zhou, Zhongyi Ye 等EMNLP 2023 · 被引用 1 次
它引用的顶会 Paper8
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- CM-BERT: Cross-Modal BERT for Text-Audio Sentiment AnalysisKaicheng Yang, Hua Xu, Kai GaoACM MM 2020 · 被引用 129 次
- Learning Hierarchical Discrete Linguistic Units from Visually-Grounded SpeechDavid Harwath, Wei-Ning Hsu, James R. GlassICLR 2020 · 被引用 88 次
- Making Monolingual Sentence Embeddings Multilingual using Knowledge DistillationNils Reimers, Iryna GurevychEMNLP 2020 · 被引用 54 次
- Language-agnostic BERT Sentence EmbeddingFangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan 等ACL 2022
相关 Paper
- Simple and Effective Unsupervised Speech TranslationChanghan Wang, Hirofumi Inaguma, Peng-Jen Chen, Ilia Kulikov 等ACL 2023 · 被引用 9 次
- CCMatrix: Mining Billions of High-Quality Parallel Sentences on the WebHolger Schwenk, Guillaume Wenzek, Sergey Edunov, Edouard Grave 等ACL 2021
- Improving End-to-End Speech Translation by Leveraging Auxiliary Speech and Text DataYuhao Zhang, Chen Xu, Bojie Hu, Chunliang Zhang 等AAAI 2023 · 被引用 17 次
- Speech Vecalign: an Embedding-based Method for Aligning Parallel Speech DocumentsChutong Meng, Philipp KoehnEMNLP 2025
- Consecutive Decoding for Speech-to-text TranslationQianqian Dong, Mingxuan Wang, Hao Zhou, Shuang Xu 等AAAI 2021 · 被引用 46 次
