Multimodal Latent Language Modeling with Next-Token Diffusion
Yutao Sun, Hangbo Bao, Wenhui Wang, Zhiliang Peng, Li Dong, Shaohan Huang, Yaoyao Chang, Jianyong Wang, Furu Wei
Abstract
Multimodal generative models require a unified approach to handle both discrete data (e.g., text and code) and continuous data (e.g., image, audio, video). In this work, we propose Latent Language Modeling (LatentLM), which seamlessly integrates continuous and discrete data using causal Transformers. Specifically, we employ a variational autoencoder (VAE) to represent continuous data as latent vectors and introduce next-token diffusion for autoregressive generation of these vectors. Additionally, we develop σ-VAE to address the challenges of variance collapse, which is crucial for autoregressive modeling. Extensive experiments demonstrate the effectiveness of LatentLM across various modalities. In image generation, LatentLM surpasses Diffusion Transformers in both performance and scalability. When integrated into multimodal large language models, LatentLM provides a general-purpose interface that unifies multimodal generation and understanding. Experimental results show that LatentLM achieves favorable performance compared to Transfusion and vector quantized models in the setting of scaling up training tokens. In text-to-speech synthesis, LatentLM outperforms the state-ofthe-art VALL-E 2 model in speaker similarity and robustness, while requiring 10× fewer decoding steps. The results establish LatentLM as a highly effective and scalable approach to advance large multimodal models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 339af1ba-e2c1-40f9-b0ee-efedc28d478cCited by top-tier papers19
- WISE: World Knowledge-Informed Semantic Evaluation for Text-to-Image GenerationYuwei Niu, Munan Ning, Mengren Zheng, Weiyang Jin et al.ICML 2026 · 195 citations
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at ScaleChunrui Han, Guopeng Li, Jingwei Wu, Quan Sun et al.ICLR 2026 · 58 citations
- SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion RefinementChenyu Yang, Shuai Wang, Hangting Chen, Wei Tan et al.NeurIPS 2025 · 28 citations
- TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language ModelingLiang-Hsuan Tseng, Yi-Chang Chen, Kuan Yi Lee, Da-shan Shiu et al.ICLR 2026 · 26 citations
- Hyperspherical Latents Improve Continuous-Token Autoregressive GenerationGuolin Ke, Hui XueICLR 2026 · 19 citations
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
Related papers
- Latent Diffusion for Language GenerationJustin Lovelace, Varsha Kishore, Chao Wan, Eliot Shekhtman et al.NeurIPS 2023 · 177 citations
- KALL-E: Autoregressive Speech Synthesis with Next-Distribution PredictionKangxiang Xia, Xinfa Zhu, Jixun Yao, Wenjie Tian et al.AAAI 2026 · 3 citations
- Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token PredictionShu-Wen Yang, Byeonggeun Kim, Kuan-Po Huang, Qingming Tang et al.ICML 2025
- Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech SynthesisWeiwei Lin, Chenhang HeICLR 2025
- Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal ModelChunting Zhou, Lili Yu, Arun Babu, Kushal Tirumala et al.ICLR 2025
