Generative Pre-trained Autoregressive Diffusion Transformer
Yuan Zhang, Jiacheng Jiang, Guoqing Ma, Zhiying Lu, Bo Wang, Haoyang Huang, Jianlong Yuan, Nan Duan
Abstract
In this work, we present GPDiT, a Generative Pre-trained Autoregressive Diffusion Transformer that unifies the strengths of diffusion and autoregressive modeling for long-range video synthesis, within a continuous latent space. Instead of predicting discrete tokens, GPDiT autoregressively predicts future latent frames using a diffusion loss, enabling natural modeling of motion dynamics and semantic consistency across frames. This continuous autoregressive framework not only enhances generation quality but also endows the model with representation capabilities. Additionally, we introduce a lightweight causal attention variant and a parameter-free rotation-based time-conditioning mechanism, improving both the training and inference efficiency. Extensive experiments demonstrate that GPDiT achieves strong performance in video generation quality, video representation ability, and few-shot learning tasks, highlighting its potential as an effective framework for video modeling in continuous space.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video DiffusionXun Huang, Zhengqi Li, Guande He, Mingyuan Zhou et al.NeurIPS 2025 · 628 citations
- BAgger: Backwards Aggregation for Mitigating Drift in Autoregressive Video Diffusion ModelsRyan Po, Eric Ryan Chan, Changan Chen, Gordon WetzsteinCVPR 2026 · 18 citations
- Streaming Autoregressive Video Generation via Diagonal DistillationJinxiu Liu, Xuanming Liu, Kangfu Mei, Yandong Wen et al.ICLR 2026 · 16 citations
- Towards Holistic Modeling for Video Frame Interpolation with Auto-regressive Diffusion TransformersXinyu Peng, Han Li, Yuyang Huang, Ziyang Zheng et al.CVPR 2026 · 4 citations
- HL-OutPaint: Coarse-to-Fine Video Outpainting for High-Resolution Long-Range VideosJeongeun Park, Janghyeok Han, Geonung Kim, Hyun-Seung Lee et al.SIGGRAPH 2026
Builds on28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
Related papers
- VDT: General-purpose Video Diffusion Transformers via Mask ModelingHaoyu Lu, Guoxing Yang, Nanyi Fei, Yuqi Huo et al.ICLR 2024 · 117 citations
- REGEN: Learning Compact Video Embedding with (Re-)Generative DecoderYitian Zhang, Long Mai, Aniruddha Mahapatra, David Bourgin et al.ICCV 2025
- PanoDiT: Panoramic Videos Generation with Diffusion TransformerMuyang Zhang, Yuzhi Chen, Rongtao Xu, Changwei Wang et al.AAAI 2025 · 6 citations
- Autoregressive Video Generation without Vector QuantizationHaoge Deng, Ting Pan, Haiwen Diao, Zhengxiong Luo et al.ICLR 2025
- AV-DiT: Taming Image Diffusion Transformers for Efficient Joint Audio and Video GenerationKai Wang, Shijian Deng, Jing Shi, Dimitrios Hatzinakos et al.ACM MM 2025 · 2 citations
