SANA: Efficient High-Resolution Text-to-Image Synthesis with Linear Diffusion Transformers
Enze Xie, Junsong Chen, Junyu Chen, Han Cai, Haotian Tang, Yujun Lin, Zhekai Zhang, Muyang Li, Ligeng Zhu, Yao Lu, Song Han
Abstract
We introduce Sana, a text-to-image framework that can efficiently generate images up to 40964096 resolution. Sana can synthesize high-resolution, high-quality images with strong text-image alignment at a remarkably fast speed, deployable on laptop GPU. Core designs include: (1) Deep compression autoencoder: unlike traditional AEs, which compress images only 8, we trained an AE that can compress images 32, effectively reducing the number of latent tokens. (2) Linear DiT: we replace all vanilla attention in DiT with linear attention, which is more efficient at high resolutions without sacrificing quality. (3) Decoder-only text encoder: we replaced T5 with modern decoder-only small LLM as the text encoder and designed complex human instruction with in-context learning to enhance the image-text alignment. (4) Efficient training and sampling: we propose Flow-DPM-Solver to reduce sampling steps, with efficient caption labeling and selection to accelerate convergence. As a result, Sana-0.6B is very competitive with modern giant diffusion model (e.g. Flux-12B), being 20 times smaller and 100+ times faster in measured throughput. Moreover, Sana-0.6B can be deployed on a 16GB laptop GPU, taking less than 1 second to generate a 10241024 resolution image. Sana enables content creation at low cost. Code and model will be publicly released upon publication.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f42675f4-27aa-4f0b-a8d5-1f37a70cee60Cited by top-tier papers18
- dKV-Cache: The Cache for Diffusion Language ModelsXinyin Ma, Runpeng Yu, Gongfan Fang, Xinchao WangNeurIPS 2025 · 145 citations
- SANA-Video: Efficient Video Generation with Block Linear Diffusion TransformerJunsong Chen, Yuyang Zhao, Jincheng Yu, Ruihang Chu et al.ICLR 2026 · 96 citations
- Lost in Latent Space: An Empirical Study of Latent Diffusion Models for Physics EmulationFrançois Rozet, Ruben Ohana, Michael McCabe, Gilles Louppe et al.NeurIPS 2025 · 23 citations
- QVGen: Pushing the Limit of Quantized Video Generative ModelsYushi Huang, Ruihao Gong, Jing Liu, Yifu Ding et al.ICLR 2026 · 22 citations
- IntrinsiX: High-Quality PBR Generation using Image PriorsPeter Kocsis, Lukas Höllein, Matthias NießnerNeurIPS 2025 · 20 citations
Builds on30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
Related papers
- SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion TransformerEnze Xie, Junsong Chen, Yuyang Zhao, Jincheng Yu et al.ICML 2025
- SANA-Sprint: One-Step Diffusion with Continuous-Time Consistency DistillationJunsong Chen, Shuchen Xue, Yuyang Zhao, Jincheng Yu et al.ICCV 2025 · 6 citations
- DiT-IC: Aligned Diffusion Transformer for Efficient Image CompressionJunqi Shi, Ming Lu, Xingchen Li, Anle Ke et al.CVPR 2026 · 4 citations
- On the Scalability of Diffusion-based Text-to-Image GenerationHao Li, Yang Zou, Ying Wang, Orchid Majumder et al.CVPR 2024
- Rethinking Cross-Modal Interaction in Multimodal Diffusion TransformersZhengyao Lv, Tianlin Pan, Chenyang Si, Zhaoxi Chen et al.ICCV 2025 · 3 citations
