Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?
Qingkai Fang, Shaolei Zhang, Zhengrui Ma, Min Zhang, Yang Feng
摘要
Two-pass direct speech-to-speech translation (S2ST) models have shown promising results which decompose S2ST into speech-to-text translation (S2TT) and text-to-speech (TTS), yet conduct end-to-end training by sharing the target text representation between S2TT and TTS models. However, the training of these models still requires large-scale parallel speech data comprising <source speech, target text, target speech> triplets, which is extremely challenging to collect. On the other hand, S2TT and TTS have accumulated a large amount of data and numerous pretrained models, which can be used to reduce the reliance on parallel speech data. To this end, we propose a composite S2ST model named ComSpeech, which connects pretrained S2TT and TTS models by introducing a vocabulary adaptor based on connectionist temporal classification (CTC). The vocabulary adaptor is employed to adapt the output text sequence of S2TT to the input text sequence of TTS, which are different due to the use of different vocabularies. In this way, ComSpeech can still be trained end-to-end and only needs a small amount of parallel speech data to finetune. We further propose a novel training method ComSpeech-ZS to eliminate the reliance on parallel speech data by aligning the text representation space of S2TT and TTS. Experimental results on the CVSS dataset show that when the parallel speech data is available, ComSpeech surpasses previous two-pass models like UnitY and Translatotron 2 in both translation quality and decoding speed. When there is no parallel speech data, ComSpeech-ZS lags behind ComSpeech by only 0.7 ASR-BLEU and outperforms the cascaded models. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper16
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin 等ICLR 2021 · 被引用 513 次
- Direct Speech-to-Speech Translation With Discrete UnitsAnn Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu 等ACL 2022 · 被引用 235 次
- UWSpeech: Speech to Speech Translation for Unwritten LanguagesChen Zhang, Xu Tan, Yi Ren, Tao Qin 等AAAI 2021 · 被引用 69 次
相关 Paper
- DASpeech: Directed Acyclic Transformer for Fast and High-quality Speech-to-Speech TranslationQingkai Fang, Yan Zhou, Yang FengNeurIPS 2023 · 被引用 22 次
- UnitY: Two-pass Direct Speech-to-speech Translation with Discrete UnitsHirofumi Inaguma, Sravya Popuri, Ilia Kulikov, Peng-Jen Chen 等ACL 2023 · 被引用 30 次
- ComSL: A Composite Speech-Language Model for End-to-End Speech-to-Text TranslationChenyang Le, Yao Qian, Long Zhou, Shujie Liu 等NeurIPS 2023 · 被引用 21 次
- StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task LearningShaolei Zhang, Qingkai Fang, Shoutao Guo, Zhengrui Ma 等ACL 2024
- Pre-training for Speech Translation: CTC Meets Optimal TransportPhuong-Hang Le, Hongyu Gong, Changhan Wang, Juan Pino 等ICML 2023 · 被引用 33 次
