NUWA-Infinity: Autoregressive over Autoregressive Generation for Infinite Visual Synthesis
Jian Liang, Chenfei Wu, Xiaowei Hu, Zhe Gan, Jianfeng Wang, Lijuan Wang, Zicheng Liu, Yuejian Fang, Nan Duan
Abstract
In this paper, we present NUWA-Infinity, a generative model for infinite visual synthesis, which is defined as the task of generating arbitrarily-sized high-resolution images or long-duration videos. An autoregressive over autoregressive generation mechanism is proposed to deal with this variable-size generation task, where a global patch-level autoregressive model considers the dependencies between patches, and a local token-level autoregressive model considers dependencies between visual tokens within each patch. A Nearby Context Pool (NCP) is introduced to cache-related patches already generated as the context for the current patch being generated, which can significantly save computation costs without sacrificing patch-level dependency modeling. An Arbitrary Direction Controller (ADC) is used to decide suitable generation orders for different visual synthesis tasks and learn order-aware positional embeddings. Compared to DALL•E, Imagen and Parti, NUWA-Infinity can generate high-resolution images with arbitrary sizes and support long-duration video generation additionally. Compared to NUWA, which also covers images and videos, NUWA-Infinity has superior visual synthesis capabilities in terms of resolution and variable-size generation. The GitHub link is https://github.com/microsoft/NUWA . The homepage link is https://nuwa-infinity.microsoft.com .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a977edc6-ee4f-4f71-ac9f-b1ac73ae2a0dCited by top-tier papers23
- InfiniCity: Infinite-Scale City SynthesisChieh Hubert Lin, Hsin-Ying Lee, Willi Menapace, Menglei Chai et al.ICCV 2023 · 86 citations
- Spatia: Video Generation with Updatable Spatial MemoryJinjing Zhao, Fangyun Wei, Zhening Liu, Hongyang Zhang et al.CVPR 2026 · 37 citations
- FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive BenchmarkRongyao Fang, Aldrich Yu, Chengqi Duan, Linjiang Huang et al.ICLR 2026 · 37 citations
- Qwen-Image-Layered: Towards Inherent Editability via Layer DecompositionShengming Yin, Zekai Zhang, Zecheng Tang, Kaiyuan Gao et al.CVPR 2026 · 30 citations
- GlueGen: Plug and Play Multi-modal Encoders for X-to-image GenerationCan Qin, Ning Yu, Chen Xing, Shu Zhang et al.ICCV 2023 · 27 citations
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- NUWA-XL: Diffusion over Diffusion for eXtremely Long Video GenerationShengming Yin, Chenfei Wu, Huan Yang, Jianfeng Wang et al.ACL 2023 · 40 citations
- InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual GenerationJinlai Liu, Jian Han, Bin Yan, Hui Wu et al.NeurIPS 2025 · 45 citations
- D-AR: Diffusion via Autoregressive ModelsZiteng Gao, Mike Zheng ShouICLR 2026 · 11 citations
- InfinityGAN: Towards Infinite-Pixel Image SynthesisChieh Hubert Lin, Hsin-Ying Lee, Yen-Chi Cheng, Sergey Tulyakov et al.ICLR 2022 · 84 citations
- Endless World: Real-Time 3D-Aware Long Video GenerationKe Zhang, Jiacong Xu, Yiqun Mei, Vishal M. PatelCVPR 2026 · 4 citations
