MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation
Khai Le-Duc, Tuyen Tran, Bach Phan Tat, Nguyen Kim Hai Bui, Quan Dang Anh, Hung-Phong Tran, Thanh Thuy Nguyen, Ly Nguyen, Tuan-Minh Phan, Thi Thu Phuong Tran, Chris Ngo, Nguyen X. Khanh, Thanh Nguyen-Tang
摘要
Multilingual speech translation (ST) and machine translation (MT) in the medical domain enhances patient care by enabling efficient communication across language barriers, alleviating specialized workforce shortages, and facilitating improved diagnosis and treatment, particularly during pandemics. In this work, we present the first systematic study on medical ST, to our best knowledge, by releasing MultiMed-ST , a large-scale ST dataset for the medical domain, spanning all translation directions in five languages: Vietnamese, English, German, French, and Simplified/Traditional Chinese, together with the models. With 290,000 samples, this is the largest medical MT dataset and the largest many-tomany multilingual ST among all domains. Secondly, we present the most comprehensive ST analysis in the field's history, to our best knowledge, including: empirical baselines, bilingual-multilingual comparative study, endto-end vs. cascaded comparative study, taskspecific vs. multi-task sequence-to-sequence comparative study, code-switch analysis, and quantitative-qualitative error analysis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
相关 Paper
- MedDialog: Large-scale Medical Dialogue DatasetsGuangtao Zeng, Wenmian Yang, Zeqian Ju, Yue Yang 等EMNLP 2020 · 被引用 163 次
- CMU-MOSEAS: A Multimodal Language Dataset for Spanish, Portuguese, German and FrenchAmirAli Bagher Zadeh, Yansheng Cao, Smon Hessner, Paul Pu Liang 等EMNLP 2020 · 被引用 43 次
- DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical domainsYanis Labrak, Adrien Bazoge, Richard Dufour, Mickael Rouvier 等ACL 2023 · 被引用 19 次
- MLSUM: The Multilingual Summarization CorpusThomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski 等EMNLP 2020 · 被引用 4 次
- Towards Building Large Scale Datasets and State-of-the-Art Automatic Speech Translation Systems for 14 Indian LanguagesAshwin Sankar, Sparsh Jain, Nikhil Narasimhan, Devilal Choudhary 等ACL 2025 · 被引用 5 次
