PixelDiT: Pixel Diffusion Transformers for Image Generation
Yongsheng Yu, Wei Xiong, Weili Nie, Yichen Sheng, Shiqiu Liu, Jiebo Luo
Abstract
Latent-space modeling has been the standard for Diffusion Transformers (DiTs). However, it relies on a two-stage pipeline where the pretrained autoencoder introduces lossy reconstruction, leading to error accumulation while hindering joint optimization. To address these issues, we propose PixelDiT, a single-stage, end-to-end model that eliminates the need for the autoencoder and learns the diffusion process directly in the pixel space. PixelDiT adopts a fully transformer-based architecture shaped by a dual-level design: a patch-level DiT that captures global semantics and a pixel-level DiT that refines texture details, enabling efficient training of a pixel-space diffusion model while preserving fine details. PixelDiT achieves 1.61 FID on ImageNet 256 and 1.81 FID on ImageNet 512, surpassing existing pixel generative models. We further extend PixelDiT to text-to-image generation and pretrain it at the 1024 2 resolution in pixel space. It achieves 0.74 on GenEval and 83.5 on DPG-bench, approaching the best latent diffusion models. Links: GitHub Code | HF Models | Project Page (a) High-resolution text-to-image samples at the megapixel scale (approximately 1024×1024) generated by PixelDiT-T2I, which is directly trained on pixel space.
"A bicycle parked on the sidewalk…" → "A motorcycle parked on the sidewalk…" SD3 VAE Recon.
FLUX VAE Recon.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c00808e5-ed12-49c1-8019-ab42f82d8f2cCited by top-tier papers4
- RePack then Refine: Efficient Diffusion Transformers with Vision Foundation ModelsGuanfang Dong, Luke Schultz, Negar Hassanpour, Chao GaoICML 2026 · 1 citation
- NNiT: Width-Agnostic Neural Network Generation with Structurally Aligned Weight SpacesJiwoo Kim, Swarajh Mehta, Hao-Lun Hsu, Hyunwoo Ryu et al.ICML 2026
- One-step Latent-free Image Generation with Pixel Mean FlowsYiyang Lu, Susie Lu, Qiao Sun, Hanhong Zhao et al.ICML 2026
- Matérn Noise for Triangulation-Agnostic Flow Matching on MeshesTianshu Kuai, Arman Maesumi, Daniel Ritchie, Noam AigermanSIGGRAPH 2026
Builds on30
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image SynthesisJunsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao et al.ICLR 2024 · 831 citations
Related papers
- PixNerd: Pixel Neural Field DiffusionShuai Wang, Ziteng Gao, Chenhui Zhu, Weilin Huang et al.ICLR 2026 · 78 citations
- DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image GenerationZehong Ma, Longhui Wei, Shuai Wang, Shiliang Zhang et al.CVPR 2026 · 59 citations
- DiP: Taming Diffusion Models in Pixel SpaceZhennan Chen, Junwei Zhu, Xu Chen, Jiangning Zhang et al.CVPR 2026 · 46 citations
- There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-TrainingJiachen Lei, Keli Liu, Julius Berner, Y HoiM et al.ICLR 2026 · 24 citations
- SpeeDiff: Scalable Pixel-Anchored End-to-End Latent Diffusion ModelBingliang Zhang, Wenda Chu, Yizhuo Li, Linjie Yang et al.CVPR 2026
