Jointly Training Large Autoregressive Multimodal Models
Emanuele Aiello, Lili Yu, Yixin Nie, Armen Aghajanyan, Barlas Oguz
Abstract
In recent years, advances in the large-scale pretraining of language and text-to-image models have revolutionized the field of machine learning. Yet, integrating these two modalities into a single, robust model capable of generating seamless multimodal outputs remains a significant challenge. To address this gap, we present the Joint Autoregressive Mixture (JAM) framework, a modular approach that systematically fuses existing text and image generation models. We also introduce a specialized, data-efficient instruction-tuning strategy, tailored for mixed-modal generation tasks. Our final instruct-tuned model demonstrates unparalleled performance in generating high-quality multimodal outputs and represents the first model explicitly designed for this purpose.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5a5aa67e-13a2-4c52-807f-3749575cb235Cited by top-tier papers13
- MMaDA: Multimodal Large Diffusion Language ModelsLing Yang, Ye Tian, Bowen Li, Xinchen Zhang et al.NeurIPS 2025 · 255 citations
- Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and ActionJiasen Lu, Christopher Clark, Sangho Lee, Zichen Zhang et al.CVPR 2024 · 53 citations
- Grounding Multimodal Large Language Models in ActionsAndrew Szot, Bogdan Mazoure, Harsh Agrawal, R. Devon Hjelm et al.NeurIPS 2024 · 43 citations
- World Model on Million-Length Video And Language With Blockwise RingAttentionHao Liu, Wilson Yan, Matei Zaharia, Pieter AbbeelICLR 2025 · 11 citations
- When Shared Knowledge Hurts: Spectral Over-Accumulation in Model MergingYayuan Li, Ze Peng, Jian Zhang, Jintao Guo et al.ICML 2026 · 5 citations
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and GenerationChengyue Wu, Xiaokang Chen, Zhiyu Wu, Yiyang Ma et al.CVPR 2025
- X-Fusion: Introducing New Modality to Frozen Large Language ModelsSicheng Mo, Thao Nguyen, Xun Huang, Siddharth Srinivasan Iyer et al.ICCV 2025
- HaploVL: A Single-Transformer Baseline for Multi-Modal UnderstandingRui Yang, Lin Song, Yicheng Xiao, Runhui Huang et al.ICML 2025
- LMFusion: Adapting Pretrained Language Models for Multimodal GenerationWeijia Shi, Xiaochuang Han, Chunting Zhou, Weixin Liang et al.NeurIPS 2025 · 134 citations
- Instruct-Imagen: Image Generation with Multi-modal InstructionHexiang Hu, Kelvin C. K. Chan, Yu-Chuan Su, Wenhu Chen et al.CVPR 2024
