Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model Perspective
Hangjie Yuan, Weihua Chen, Jun Cen, Hu Yu, Jingyun Liang, Shuning Chang, Zhihui Lin, Tao Feng, Pengwei Liu, Jiazheng Xing, Hao Luo, Jiasheng Tang
Abstract
Autoregressive large language models (LLMs) have unified a vast range of language tasks, inspiring preliminary efforts in autoregressive (AR) video generation. Existing AR video generators either diverge from standard LLM architectures, depend on bulky external text encoders, or incur prohibitive latency due to next-token decoding. In this paper, we introduce Lumos-1, an LLM-based unified model for AR video generation with efficient discrete diffusion. Firstly, to fit videos with LLMs, we identify that 1D RoPE is ill-suited for visual spatiotemporal correlation modeling, and while demonstrated to be useful, naive 3D RoPE exhibits imbalanced frequency spectra. Therefore, we propose MM-RoPE, which preserves the original textual RoPE while seamlessly accommodating video data with comprehensive frequency spectra and scaled 3D positions. Secondly, to fit the video data's nature and overcome the inefficiency of next-token decoding, we adopt a parallel and mask-based discrete diffusion with the intra-frame bidirectional and inter-frame causal attention masks. Based on this attention mask, we uncover the frame-wise loss imbalance issue caused by spatial information redundancy and propose Autoregressive Discrete Diffusion Forcing, which introduces temporal tube masking during training with a compatible inference-time masking policy to avoid quality degradation. Despite using only 48 GPUs for pre-training and fine-tuning, limited data and a discrete tokenizer, Lumos-1 achieves results surpassing those of Show-o2 on GenEval, COSMOS-Video2World on VBench-I2V, and OpenSoraPlan on VBench-T2V. Code and models are available at https://github.com/alibaba-damo-academy/Lumos.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0f5d130c-2dd4-477d-a336-a2fa26c1e623Cited by top-tier papers9
- Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching DistillationYunhong Lu, Yanhong Zeng, Haobo Li, Hao Ouyang et al.CVPR 2026 · 77 citations
- VideoWorld 2: Learning Transferable Knowledge from Real-world VideosZhongwei Ren, Yunchao Wei, Xiao Yu, Guixun Luo et al.CVPR 2026 · 9 citations
- MaskFocus: Focusing Policy Optimization on Critical Steps for Masked Image GenerationGuohui Zhang, Hu Yu, Xiaoxiao Ma, Yaning Pan et al.CVPR 2026 · 6 citations
- TV2TV: A Unified Framework for Interleaved Language and Video GenerationXiaochuang Han, Youssef Emad, Melissa Hall, John Nguyen et al.CVPR 2026 · 3 citations
- OmniLottie: Generating Vector Animations via Parameterized Lottie TokensYiying Yang, Wei Cheng, Sijin Chen, Honghao Fu et al.CVPR 2026 · 2 citations
Builds on42
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- LumosX: Relate Any Identities with Their Attributes for Personalized Video GenerationJiazheng Xing, Fei Du, Hangjie Yuan, Pengwei Liu et al.ICLR 2026 · 6 citations
- Infinity-RoPE: Action-Controllable Infinite Video Generation Emerges From Autoregressive Self-RolloutHidir Yesiltepe, Tuna Han Salih Meral, Adil Kaan Akan, Kaan Oktay et al.CVPR 2026 · 84 citations
- VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video UnderstandingZhihao He, Tieyuan Chen, Kangyu Wang, Ziran Qin et al.ICML 2026 · 3 citations
- Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language ModelsSujung Hong, Chanyong Yoon, Seong Jae HwangICML 2026
- UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible FeedbackPengwei Liu, Hangjie Yuan, Bo Dong, Jiazheng Xing et al.NeurIPS 2025 · 10 citations
