VideoMAR: Autoregressive Video Generation with Continuous Tokens
Hu Yu, Biao Gong, Hangjie Yuan, Dandan Zheng, Weilong Chai, Jingdong Chen, Kecheng Zheng, Feng Zhao
摘要
Mask-based autoregressive models have demonstrated promising image generation capability in continuous space. However, their potential for video generation remains under-explored. In this paper, we propose VideoMAR, a concise and efficient decoder-only autoregressive image-to-video model with continuous tokens, composing temporal frame-by-frame and spatial masked generation. We first identify temporal causality and spatial bi-directionality as the first principle of video AR models, and propose the next-frame diffusion loss for the integration of mask and video generation. Besides, the huge cost and difficulty of long sequence autoregressive modeling is a basic but crucial issue. To this end, we propose the temporal short-to-long curriculum learning and spatial progressive resolution training, and employ progressive temperature strategy at inference time to mitigate the accumulation error. Furthermore, VideoMAR replicates several unique capacities of language models to video generation. It inherently bears high efficiency due to simultaneous temporal-wise KV cache and spatial-wise parallel generation, and presents the capacity of spatial and temporal extrapolation via * Project lead † Corresponding author 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
3D rotary embeddings. On the VBench-I2V benchmark, VideoMAR surpasses the previous state-of-the-art (Cosmos I2V) while requiring significantly fewer parameters (9.3%), training data (0.5%), and GPU resources (0.2%).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model PerspectiveHangjie Yuan, Weihua Chen, Jun Cen, Hu Yu 等ICLR 2026 · 被引用 21 次
- Causality in Video Diffusers is Separable from DenoisingXingjian Bai, Guande He, Zhengqi Li, Eli Shechtman 等CVPR 2026 · 被引用 3 次
它引用的顶会 Paper17
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Autoregressive Image Generation without Vector QuantizationTianhong Li, Yonglong Tian, He Li, Mingyang Deng 等NeurIPS 2024 · 被引用 758 次
- Diffusion Forcing: Next-token Prediction Meets Full-Sequence DiffusionBoyuan Chen, Diego Marti Monso, Yilun Du, Max Simchowitz 等NeurIPS 2024 · 被引用 751 次
- Genie: Generative Interactive EnvironmentsJake Bruce, Michael D. Dennis, Ashley Edwards, Jack Parker-Holder 等ICML 2024 · 被引用 513 次
相关 Paper
- Autoregressive Video Generation without Vector QuantizationHaoge Deng, Ting Pan, Haiwen Diao, Zhengxiong Luo 等ICLR 2025
- Phenaki: Variable Length Video Generation from Open Domain Textual DescriptionsRuben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans, Hernan Moraldo 等ICLR 2023 · 被引用 79 次
- InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual GenerationJinlai Liu, Jian Han, Bin Yan, Hui Wu 等NeurIPS 2025 · 被引用 45 次
- High-Quality Joint Image and Video Tokenization with Causal VAEDawit Mureja Argaw, Xian Liu, Qinsheng Zhang, Joon Son Chung 等ICLR 2025
- Make-A-Video: Text-to-Video Generation without Text-Video DataUriel Singer, Adam Polyak, Thomas Hayes, Xi Yin 等ICLR 2023 · 被引用 313 次
