Revisiting Non-Autoregressive Transformers for Efficient Image Synthesis
Zanlin Ni, Yulin Wang, Renping Zhou, Jiayi Guo, Jinyi Hu, Zhiyuan Liu, Shiji Song, Yuan Yao, Gao Huang
Abstract
The field of image synthesis is currently flourishing due to the advancements in diffusion models. While diffusion models have been successful, their computational inten-sity has prompted the pursuit of more efficient alternatives. As a representative work, non-autoregressive Transformers (NATs) have been recognized for their rapid generation. However, a major drawback of these models is their in-ferior performance compared to diffusion models. In this paper, we aim to re-evaluate the full potential of NATs by revisiting the design of their training and inference strategies. Specifically, we identify the complexities in properly configuring these strategies and indicate the possible sub-optimality in existing heuristic-driven designs. Recognizing this, we propose to go beyond existing methods by directly solving the optimal strategies in an automatic framework. The resulting method, named AutoNAT, advances the performance boundaries of NATs notably, and is able to perform comparably with the latest diffusion models with a significantly reduced inference cost. The effectiveness of AutoNAT is comprehensively validated on four benchmark datasets, i.e., ImageNet-256 & 512, MS-COCO, and CC3M. Code and pretrained models will be available at htt P s: / /gi thub. com/LeapLabTHU/ImprovedNAT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 66820d75-f73a-4d33-8c07-5ba2cf35f600Cited by top-tier papers17
- Bridging the Divide: Reconsidering Softmax and Linear AttentionDongchen Han, Yifan Pu, Zhuofan Xia, Yizeng Han et al.NeurIPS 2024 · 69 citations
- Real-world Image Dehazing with Coherence-based Pseudo Labeling and Cooperative Unfolding NetworkChengyu Fang, Chunming He, Fengyang Xiao, Yulun Zhang et al.NeurIPS 2024 · 46 citations
- GSVA: Generalized Segmentation via Multimodal Large Language ModelsZhuofan Xia, Dongchen Han, Yizeng Han, Xuran Pan et al.CVPR 2024 · 42 citations
- ENAT: Rethinking Spatial-temporal Interactions in Token-based Image SynthesisZanlin Ni, Yulin Wang, Renping Zhou, Yizeng Han et al.NeurIPS 2024 · 15 citations
- CODA: Repurposing Continuous VAEs for Discrete TokenizationZeyu Liu, Zanlin Ni, Yeguo Hua, Xin Deng et al.ICCV 2025 · 9 citations
Builds on21
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Denoising Autoregressive Transformers for Scalable Text-to-Image GenerationJiatao Gu, Yuyang Wang, Yizhe Zhang, Qihang Zhang et al.ICLR 2025
- RenewNAT: Renewing Potential Translation for Non-autoregressive TransformerPei Guo, Yisheng Xiao, Juntao Li, Min ZhangAAAI 2023 · 9 citations
- AEQA-NAT : Adaptive End-to-end Quantization Alignment Training Framework for Non-autoregressive Machine TranslationXiangyu Qu, Guojing Liu, Liang LiICML 2025
- Tread: Token Routing for Efficient Architecture-Agnostic Diffusion TrainingFelix Krause, Timy Phan, Ming Gui, Stefan Andreas Baumann et al.ICCV 2025 · 1 citation
- NAT4AT: Using Non-Autoregressive Translation Makes Autoregressive Translation Faster and BetterHuanran Zheng, Wei Zhu, Xiaoling WangWWW 2024 · 13 citations
