Diffusion Models Without Attention
Jing Nathan Yan, Jiatao Gu, Alexander M. Rush
Abstract
In recent advancements in high-fidelity image generation, Denoising Diffusion Probabilistic Models (DDPMs) have emerged as a key player. However, their application at high resolutions presents significant computational challenges. Current methods, such as patchifying, expedite processes in UNet and Transformer architectures but at the expense of rep-resentational capacity. Addressing this, we introduce the Dif-fusion State Space Model (DIFFUSSM), an architecture that supplants attention mechanisms with a more scalable state space model backbone. This approach effectively handles higher resolutions without resorting to global compression, thus preserving detailed image representation throughout the diffusion process. Our focus on FLOP-efficient architectures in diffusion training marks a significant step forward. Comprehensive evaluations on both ImageNet and LSUN datasets at two resolutions demonstrate that DiffuSSMs are on par or even outperform existing diffusion models with attention modules in FID and Inception Score metrics while significantly reducing total FLOP usage.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe7932d5-d67e-49b2-a56f-158049285fccCited by top-tier papers42
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- Parallelizing Linear Transformers with the Delta Rule over Sequence LengthSonglin Yang, Bailin Wang, Yu Zhang, Yikang Shen et al.NeurIPS 2024 · 412 citations
- Gated Linear Attention Transformers with Hardware-Efficient TrainingSonglin Yang, Bailin Wang, Yikang Shen, Rameswar Panda et al.ICML 2024 · 390 citations
- Demystify Mamba in Vision: A Linear Attention PerspectiveDongchen Han, Ziyi Wang, Zhuofan Xia, Yizeng Han et al.NeurIPS 2024 · 287 citations
- Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient InferenceHan Zhao, Min Zhang, Wei Zhao, Pengxiang Ding et al.AAAI 2025 · 125 citations
Builds on38
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- DiP: Taming Diffusion Models in Pixel SpaceZhennan Chen, Junwei Zhu, Xu Chen, Jiangning Zhang et al.CVPR 2026 · 46 citations
- Denoising Diffusion Step-aware ModelsShuai Yang, Yukang Chen, Luozhou Wang, Shu Liu et al.ICLR 2024 · 27 citations
- ScaleDiff: Higher-Resolution Image Synthesis via Efficient and Model-Agnostic DiffusionSungho Koh, SeungJu Cha, Hyunwoo Oh, Kwanyoung Lee et al.NeurIPS 2025 · 5 citations
- Scalable High-Resolution Pixel-Space Image Synthesis with Hourglass Diffusion TransformersKatherine Crowson, Stefan Andreas Baumann, Alex Birch, Tanishq Mathew Abraham et al.ICML 2024 · 98 citations
- UDPM: Upsampling Diffusion Probabilistic ModelsShady Abu-Hussein, Raja GiryesNeurIPS 2024 · 10 citations
