MIO: A Foundation Model on Multimodal Tokens
Zekun Moore Wang, King Zhu, Chunpu Xu, Wangchunshu Zhou, Jiaheng Liu, Yibo Zhang, Jessie Jiashuo Wang, Ning Shi, Siyu Li, Yizhi Li, Haoran Que, Zhaoxiang Zhang
Abstract
In this paper, we introduce MIO, a novel foundation model built on multimodal tokens, capable of understanding and generating speech, text, images, and videos in an end-to-end, autoregressive manner. While the emergence of large language models (LLMs) and multimodal large language models (MM-LLMs) propels advancements in artificial general intelligence through their versatile capabilities, they still lack true any-to-any understanding and generation. Recently, the release of GPT-4o has showcased the remarkable potential of any-to-any LLMs for complex real-world tasks, enabling omnidirectional input and output across images, speech, and text. However, it is closed-source and does not support the generation of multimodal interleaved sequences. To address this gap, we present MIO, which is trained on a mixture of discrete tokens across four modalities using causal multimodal modeling. MIO undergoes a four-stage training process: (1) alignment pre-training, (2) interleaved pre-training, (3) speech-enhanced pre-training, and (4) comprehensive supervised fine-tuning on diverse textual, visual, and speech tasks. Our experimental results indicate that MIO exhibits competitive, and in some cases superior, performance compared to previous dual-modal baselines, any-to-any model baselines, and even modality-specific baselines. Moreover, MIO demonstrates advanced capabilities inherent to its any-to-any feature, such as interleaved video-text generation, chain-of-visual-thought reasoning, visual guideline generation, instructional image editing, etc.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe29197e-18e5-4bd5-92f4-58a62f0395a4Cited by top-tier papers2
- FlowBind: Efficient Any-to-Any Generation with Bidirectional FlowsYeonwoo Cha, Semin Kim, Jinhyeon Kwon, Seunghoon HongICLR 2026 · 2 citations
- BlendServe: Optimizing Offline Inference with Resource-Aware BatchingYilong Zhao, Shuo Yang, Kan Zhu, Lianmin Zheng et al.ASPLOS 2026
Builds on37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
Related papers
- NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow MatchingRun Luo, Xiaobo Xia, Lu Wang, Longze Chen et al.ICLR 2026 · 22 citations
- EMOVA: Empowering Language Models to See, Hear and Speak with Vivid EmotionsKai Chen, Yunhao Gou, Runhui Huang, Zhili Liu et al.CVPR 2025
- CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any GenerationZineng Tang, Ziyi Yang, Mahmoud Khademi, Yang Liu et al.CVPR 2024 · 15 citations
- Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete DiffusionLijiang Li, zuwei long, Yunhang Shen, Heting Gao et al.ICML 2026 · 7 citations
- 4M-21: An Any-to-Any Vision Model for Tens of Tasks and ModalitiesRoman Bachmann, Oguzhan Fatih Kar, David Mizrahi, Ali Garjani et al.NeurIPS 2024 · 60 citations
