PixelDiT: Pixel Diffusion Transformers for Image Generation
Yongsheng Yu, Wei Xiong, Weili Nie, Yichen Sheng, Shiqiu Liu, Jiebo Luo
摘要
Latent-space modeling has been the standard for Diffusion Transformers (DiTs). However, it relies on a two-stage pipeline where the pretrained autoencoder introduces lossy reconstruction, leading to error accumulation while hindering joint optimization. To address these issues, we propose PixelDiT, a single-stage, end-to-end model that eliminates the need for the autoencoder and learns the diffusion process directly in the pixel space. PixelDiT adopts a fully transformer-based architecture shaped by a dual-level design: a patch-level DiT that captures global semantics and a pixel-level DiT that refines texture details, enabling efficient training of a pixel-space diffusion model while preserving fine details. PixelDiT achieves 1.61 FID on ImageNet 256 and 1.81 FID on ImageNet 512, surpassing existing pixel generative models. We further extend PixelDiT to text-to-image generation and pretrain it at the 1024 2 resolution in pixel space. It achieves 0.74 on GenEval and 83.5 on DPG-bench, approaching the best latent diffusion models. Links: GitHub Code | HF Models | Project Page (a) High-resolution text-to-image samples at the megapixel scale (approximately 1024×1024) generated by PixelDiT-T2I, which is directly trained on pixel space.
"A bicycle parked on the sidewalk…" → "A motorcycle parked on the sidewalk…" SD3 VAE Recon.
FLUX VAE Recon.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- RePack then Refine: Efficient Diffusion Transformers with Vision Foundation ModelsGuanfang Dong, Luke Schultz, Negar Hassanpour, Chao GaoICML 2026 · 被引用 1 次
- NNiT: Width-Agnostic Neural Network Generation with Structurally Aligned Weight SpacesJiwoo Kim, Swarajh Mehta, Hao-Lun Hsu, Hyunwoo Ryu 等ICML 2026
- One-step Latent-free Image Generation with Pixel Mean FlowsYiyang Lu, Susie Lu, Qiao Sun, Hanhong Zhao 等ICML 2026
- Matérn Noise for Triangulation-Agnostic Flow Matching on MeshesTianshu Kuai, Arman Maesumi, Daniel Ritchie, Noam AigermanSIGGRAPH 2026
它引用的顶会 Paper30
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image SynthesisJunsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao 等ICLR 2024 · 被引用 831 次
相关 Paper
- PixNerd: Pixel Neural Field DiffusionShuai Wang, Ziteng Gao, Chenhui Zhu, Weilin Huang 等ICLR 2026 · 被引用 78 次
- DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image GenerationZehong Ma, Longhui Wei, Shuai Wang, Shiliang Zhang 等CVPR 2026 · 被引用 59 次
- DiP: Taming Diffusion Models in Pixel SpaceZhennan Chen, Junwei Zhu, Xu Chen, Jiangning Zhang 等CVPR 2026 · 被引用 46 次
- There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-TrainingJiachen Lei, Keli Liu, Julius Berner, Y HoiM 等ICLR 2026 · 被引用 24 次
- SpeeDiff: Scalable Pixel-Anchored End-to-End Latent Diffusion ModelBingliang Zhang, Wenda Chu, Yizhuo Li, Linjie Yang 等CVPR 2026
