Next Patch Prediction for AutoRegressive Visual Generation
Yatian Pang, Peng Jin, Shuo Yang, Bin Zhu, Bin Lin, Chaoran Feng, Zhenyu Tang, Liuhan Chen, Francis E. H. Tay, Ser-Nam Lim, Harry Yang, Li Yuan
摘要
Autoregressive models, built based on the Next Token Prediction (NTP) paradigm, show great potential in developing a unified framework that integrates both language and vision tasks. Pioneering works introduce NTP to autoregressive visual generation tasks. In this work, we rethink the NTP for autoregressive image generation and extend it to a novel Next Patch Prediction (NPP) paradigm. Our key idea is to group and aggregate image tokens into patch tokens with higher information density. By using patch tokens as a more compact input sequence, the autoregressive model is trained to predict the next patch, significantly reducing computational costs. To further exploit the natural hierarchical structure of image data, we propose a multiscale coarse-to-fine patch grouping strategy. With this strategy, the training process begins with a large patch size and ends with vanilla NTP where the patch size is 1×1, thus maintaining the original inference process without modifications. Extensive experiments across a diverse range of model sizes demonstrate that NPP could reduce the training cost to ∼ 0.6× while improving image generation quality by up to 1.0 FID score on the ImageNet 256×256 generation benchmark. Notably, our method retains the original autoregressive model architecture without introducing additional trainable parameters or specifically designing a custom image tokenizer, offering a flexible and plug-and-play solution for enhancing autoregressive visual generation. https://github.com/PKU-YuanGroup/Next-Patch-Prediction
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Look-Back: Implicit Visual Re-focusing in MLLM ReasoningShuo Yang, Yuwei Niu, Yuyang Liu, Yang Ye 等AAAI 2026 · 被引用 32 次
- Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model PerspectiveHangjie Yuan, Weihua Chen, Jun Cen, Hu Yu 等ICLR 2026 · 被引用 21 次
- AE-NeRF: Augmenting Event-Based Neural Radiance Fields for Non-ideal Conditions and Larger ScenesChaoran Feng, Wangbo Yu, Xinhua Cheng, Zhenyu Tang 等AAAI 2025 · 被引用 21 次
- Understand Before You Generate: Self-Guided Training for Autoregressive Image GenerationXiaoyu Yue, Zidong Wang, Yuqing Wang, Wenlong Zhang 等NeurIPS 2025 · 被引用 9 次
- MVAR: Visual Autoregressive Modeling with Scale and Spatial Markovian ConditioningJinhua Zhang, Wei Long, Minghao Han, Weiyi You 等ICLR 2026 · 被引用 9 次
它引用的顶会 Paper42
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual GenerationDivyansh Srivastava, Akshay Mehra, Pranav Maneriker, Debopam Sanyal 等CVPR 2026 · 被引用 1 次
- NFIG: Multi-Scale Autoregressive Image Generation via Frequency OrderingZhihao Huang, Xi Qiu, Yukuo Ma, Yifu Zhou 等NeurIPS 2025 · 被引用 20 次
- Holistic Tokenizer for Autoregressive Image GenerationAnlin Zheng, Haochen Wang, Yucheng Zhao, Weipeng Deng 等ICCV 2025 · 被引用 11 次
- Locality-aware Parallel Decoding for Efficient Autoregressive Image GenerationZhuoyang Zhang, Luke J. Huang, Chengyue Wu, Shang Yang 等ICLR 2026 · 被引用 8 次
- Denoising Token Prediction in Masked Autoregressive ModelsTing Yao, Yehao Li, Yingwei Pan, Zhaofan Qiu 等ICCV 2025 · 被引用 2 次
