LAMA-UT: Language Agnostic Multilingual ASR Through Orthography Unification and Language-Specific Transliteration
Sangmin Lee, Woo-Jin Chung, Hong-Goo Kang
Abstract
Building a universal multilingual automatic speech recognition (ASR) model that performs equitably across languages has long been a challenge due to its inherent difficulties. To address this task we introduce a Language-Agnostic Multilingual ASR pipeline through orthography Unification and language-specific Transliteration (LAMA-UT 1 ). LAMA-UT operates without any language-specific modules while matching the performance of state-of-the-art models trained on a minimal amount of data. Our pipeline consists of two key steps. First, we utilize a universal transcription generator to unify orthographic features into Romanized form and capture common phonetic characteristics across diverse languages. Second, we utilize a universal converter to transform these universal transcriptions into language-specific ones. In experiments, we demonstrate the effectiveness of our proposed method leveraging universal transcriptions for massively multilingual ASR. Our pipeline achieves a relative error reduction rate of 45% when compared to Whisper and performs comparably to MMS, despite being trained on only 0.1% of Whisper's training data. Furthermore, our pipeline does not rely on any language-specific modules. However, it performs on par with zero-shot ASR approaches which utilize additional language-specific lexicons and language models. We expect this framework to serve as a cornerstone for flexible multilingual ASR systems that are generalizable even to unseen languages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ea39f67c-ca53-4f18-a2fc-43904a9c4f25Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- SALMONN: Towards Generic Hearing Abilities for Large Language ModelsChangli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen et al.ICLR 2024 · 557 citations
- Flashlight: Enabling Innovation in Tools for Machine LearningJacob D. Kahn, Vineel Pratap, Tatiana Likhomanenko, Qiantong Xu et al.ICML 2022 · 41 citations
Related papers
- Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech RepresentationsJeong Hun Yeo, Minsu Kim, Chae Won Kim, Stavros Petridis et al.ICCV 2025 · 3 citations
- From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech RecognitionTianduo Wang, Lu Xu, Wei Lu, Shanbo ChengEMNLP 2025 · 1 citation
- Whisper-UT: A Unified Translation Framework for Speech and TextCihan Xiao, Matthew Wiesner, Debashish Chakraborty, Reno Kriz et al.EMNLP 2025
- Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation ModelsYuchen Hu, Chen Chen, Chao-Han Huck Yang, Chengwei Qin et al.NeurIPS 2024 · 14 citations
- POWSM: A Phonetic Open Whisper-Style Speech Foundation ModelChin-Jou Li, Kalvin Chang, Shikhar Bharadwaj, Eunjung Yeo et al.ACL 2026 · 11 citations
