Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model Perspective
Hangjie Yuan, Weihua Chen, Jun Cen, Hu Yu, Jingyun Liang, Shuning Chang, Zhihui Lin, Tao Feng, Pengwei Liu, Jiazheng Xing, Hao Luo, Jiasheng Tang
摘要
Autoregressive large language models (LLMs) have unified a vast range of language tasks, inspiring preliminary efforts in autoregressive (AR) video generation. Existing AR video generators either diverge from standard LLM architectures, depend on bulky external text encoders, or incur prohibitive latency due to next-token decoding. In this paper, we introduce Lumos-1, an LLM-based unified model for AR video generation with efficient discrete diffusion. Firstly, to fit videos with LLMs, we identify that 1D RoPE is ill-suited for visual spatiotemporal correlation modeling, and while demonstrated to be useful, naive 3D RoPE exhibits imbalanced frequency spectra. Therefore, we propose MM-RoPE, which preserves the original textual RoPE while seamlessly accommodating video data with comprehensive frequency spectra and scaled 3D positions. Secondly, to fit the video data's nature and overcome the inefficiency of next-token decoding, we adopt a parallel and mask-based discrete diffusion with the intra-frame bidirectional and inter-frame causal attention masks. Based on this attention mask, we uncover the frame-wise loss imbalance issue caused by spatial information redundancy and propose Autoregressive Discrete Diffusion Forcing, which introduces temporal tube masking during training with a compatible inference-time masking policy to avoid quality degradation. Despite using only 48 GPUs for pre-training and fine-tuning, limited data and a discrete tokenizer, Lumos-1 achieves results surpassing those of Show-o2 on GenEval, COSMOS-Video2World on VBench-I2V, and OpenSoraPlan on VBench-T2V. Code and models are available at https://github.com/alibaba-damo-academy/Lumos.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching DistillationYunhong Lu, Yanhong Zeng, Haobo Li, Hao Ouyang 等CVPR 2026 · 被引用 77 次
- VideoWorld 2: Learning Transferable Knowledge from Real-world VideosZhongwei Ren, Yunchao Wei, Xiao Yu, Guixun Luo 等CVPR 2026 · 被引用 9 次
- MaskFocus: Focusing Policy Optimization on Critical Steps for Masked Image GenerationGuohui Zhang, Hu Yu, Xiaoxiao Ma, Yaning Pan 等CVPR 2026 · 被引用 6 次
- TV2TV: A Unified Framework for Interleaved Language and Video GenerationXiaochuang Han, Youssef Emad, Melissa Hall, John Nguyen 等CVPR 2026 · 被引用 3 次
- OmniLottie: Generating Vector Animations via Parameterized Lottie TokensYiying Yang, Wei Cheng, Sijin Chen, Honghao Fu 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper42
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- LumosX: Relate Any Identities with Their Attributes for Personalized Video GenerationJiazheng Xing, Fei Du, Hangjie Yuan, Pengwei Liu 等ICLR 2026 · 被引用 6 次
- Infinity-RoPE: Action-Controllable Infinite Video Generation Emerges From Autoregressive Self-RolloutHidir Yesiltepe, Tuna Han Salih Meral, Adil Kaan Akan, Kaan Oktay 等CVPR 2026 · 被引用 84 次
- VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video UnderstandingZhihao He, Tieyuan Chen, Kangyu Wang, Ziran Qin 等ICML 2026 · 被引用 3 次
- Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language ModelsSujung Hong, Chanyong Yoon, Seong Jae HwangICML 2026
- UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible FeedbackPengwei Liu, Hangjie Yuan, Bo Dong, Jiazheng Xing 等NeurIPS 2025 · 被引用 10 次
