PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Zhongdao Wang, James T. Kwok, Ping Luo, Huchuan Lu, Zhenguo Li
Abstract
The most advanced text-to-image (T2I) models require significant training costs (e.g., millions of GPU hours), seriously hindering the fundamental innovation for the AIGC community while increasing CO 2 emissions. This paper introduces PIXART-α, a Transformer-based T2I diffusion model whose image generation quality is competitive with state-of-the-art image generators (e.g., Imagen, SDXL, and even Midjourney), reaching near-commercial application standards. Additionally, it supports high-resolution image synthesis up to 1024 × 1024 resolution with low training cost, as shown in Figure 1 and 2. To achieve this goal, three core designs are proposed: (1) Training strategy decomposition: We devise three distinct training steps that respectively optimize pixel dependency, textimage alignment, and image aesthetic quality; (2) Efficient T2I Transformer: We incorporate cross-attention modules into Diffusion Transformer (DiT) to inject text conditions and streamline the computation-intensive class-condition branch; (3) High-informative data: We emphasize the significance of concept density in text-image pairs and leverage a large Vision-Language model to auto-label dense pseudo-captions to assist text-image alignment learning. As a result, PIXART-α's training speed markedly surpasses existing large-scale T2I models, e.g., PIXARTα only takes 12% of Stable Diffusion v1.5's training time (∼753 vs. ∼6,250 A100 GPU days), saving nearly 28,400 vs. $320,000) and reducing 90% CO 2 emissions. Moreover, compared with a larger SOTA model, RAPHAEL, our training cost is merely 1%. Extensive experiments demonstrate that PIXART-α excels in image quality, artistry, and semantic control. We hope PIXART-α will provide new insights to the AIGC community and startups to accelerate building their own high-quality yet low-cost generative models from scratch. INTRODUCTION Recently, the advancement of text-to-image (T2I) generative models, such as DALL•E 2 (OpenAI, 2023), Imagen (Saharia et al., 2022), and Stable Diffusion (Rombach et al., 2022) has started a new era of photorealistic image synthesis, profoundly impacting numerous downstream applications, such as image editing (Kim et al., 2022 ), video generation (Wu et al., 2022), 3D assets creation (Poole et al., 2022) , etc. However, the training of these advanced models demands immense computational resources. For instance, training SDv1.5 (Podell et al., 2023) necessitates 6K A100 GPU days, approximately costing * Equal contribution. Work done during the internships of the four students at Huawei Noah's Ark Lab.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext adf17f71-9b98-426e-b0a2-7a068cef1766Cited by top-tier papers518
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng et al.NeurIPS 2024 · 1,199 citations
- Representation Alignment for Diffusion Transformers without External ComponentsDengyang Jiang, Mengmeng Wang, Liuzhuozheng Li, Lei Zhang et al.ICLR 2026 · 532 citations
- Show-o2: Improved Native Unified Multimodal ModelsJinheng Xie, Zhenheng Yang, Mike Zheng ShouNeurIPS 2025 · 261 citations
- MMaDA: Multimodal Large Diffusion Language ModelsLing Yang, Ye Tian, Bowen Li, Xinchen Zhang et al.NeurIPS 2025 · 255 citations
- Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion TransformerShuang Wu, Youtian Lin, Yifei Zeng, Feihu Zhang et al.NeurIPS 2024 · 251 citations
Builds on36
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
Related papers
- Stretching Each Dollar: Diffusion Training from Scratch on a Micro-BudgetVikash Sehwag, Xianghao Kong, Jingtao Li, Michael Spranger et al.CVPR 2025
- KOALA: Empirical Lessons Toward Memory-Efficient and Fast Diffusion Models for Text-to-Image SynthesisYoungwan Lee, Kwanyong Park, Yoorhim Cho, Yong-Ju Lee et al.NeurIPS 2024 · 22 citations
- DiffSparse: Accelerating Diffusion Transformers with Learned Token SparsityHaowei Zhu, Ji Liu, Ziqiong Liu, Dong Li et al.ICLR 2026 · 2 citations
- Edit: Efficient Diffusion Transformers with Linear Compressed AttentionPhilipp Becker, Abhinav Mehrotra, Ruchika Chavhan, Malcolm Chadwick et al.ICCV 2025 · 9 citations
- PaGoDA: Progressive Growing of a One-Step Generator from a Low-Resolution Diffusion TeacherDongjun Kim, Chieh-Hsin Lai, Wei-Hsiang Liao, Yuhta Takida et al.NeurIPS 2024 · 32 citations
