LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Qingkai Fang, Shoutao Guo, Yan Zhou, Zhengrui Ma, Shaolei Zhang, Yang Feng
Abstract
Models like GPT-4o enable real-time interaction with large language models (LLMs) through speech, significantly enhancing user experience compared to traditional text-based interaction. However, there is still a lack of exploration on how to build speech interaction models based on open-source LLMs. To address this, we propose LLaMA-Omni, a novel end-to-end model architecture designed for low-latency and high-quality speech interaction with LLMs. LLaMA-Omni integrates a pretrained speech encoder, a speech adaptor, an LLM, and a streaming speech decoder. It eliminates the need for speech transcription, and can simultaneously generate text and speech responses directly from speech instructions with extremely low latency. We build our model based on the latest Llama-3.1-8B-Instruct model. To align the model with speech interaction scenarios, we construct a dataset named InstructS2S-200K, which includes 200K speech instructions and corresponding speech responses, with a style that better matches the characteristics of speech interaction scenarios. Experimental results show that compared to previous speech-language models, LLaMA-Omni provides better responses in both content and style, with a response latency as low as 236ms. Additionally, training LLaMA-Omni takes less than 3 days on just 4 GPUs, paving the way for the efficient development of speech-language models in the future. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09d50938-47e7-48ed-8a87-b3ee9fd6d26cCited by top-tier papers69
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech InteractionChaoyou Fu, Haojia Lin, Xiong Wang, Yifan Zhang et al.NeurIPS 2025 · 234 citations
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language ModelsWeiye Xu, Jiahao Wang, Weiyun Wang, Zhe Chen et al.ICLR 2026 · 103 citations
- OmniFlatten: An End-to-end GPT Model for Seamless Voice ConversationQinglin Zhang, Luyao Cheng, Chong Deng, Qian Chen et al.ACL 2025 · 51 citations
- SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex ConversationWenyi Yu, Siyin Wang, Xiaoyu Yang, Xianzhao Chen et al.NeurIPS 2025 · 43 citations
- STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language ModelsCheng-Han Chiang, Xiaofei Wang, Linjie Li, Chung-Ching Lin et al.ICLR 2026 · 38 citations
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 1,267 citations
- SALMONN: Towards Generic Hearing Abilities for Large Language ModelsChangli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen et al.ICLR 2024 · 557 citations
Related papers
- LLaMA-Omni 2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech SynthesisQingkai Fang, Yan Zhou, Shoutao Guo, Shaolei Zhang et al.ACL 2025
- Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLMXiong Wang, Yangze Li, Chaoyou Fu, Yike Zhang et al.ICML 2025
- Recent Advances in Speech Language Models: A SurveyWenqian Cui, Dianzhi Yu, Xiaoqi Jiao, Ziqiao Meng et al.ACL 2025
- OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-time Emotional Speech SynthesisRun Luo, Ting-En Lin, Haonan Zhang, Yuchuan Wu et al.NeurIPS 2025 · 5 citations
- OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality InteractionHaonan Zhang, Run Luo, Xiong Liu, Yuchuan Wu et al.ACL 2025 · 10 citations
