SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer
Enze Xie, Junsong Chen, Yuyang Zhao, Jincheng Yu, Ligeng Zhu, Yujun Lin, Zhekai Zhang, Muyang Li, Junyu Chen, Han Cai, Bingchen Liu, Daquan Zhou, Song Han
Abstract
This paper presents SANA-1.5, a linear Diffusion Transformer for efficient scaling in text-to-image generation. Building upon SANA-1.0, we introduce three key innovations: (1) Efficient Training Scaling: A depth-growth paradigm that enables scaling from 1.6B to 4.8B parameters with significantly reduced computational resources, combined with a memory-efficient 8-bit optimizer. (2) Model Depth Pruning: A block importance analysis technique for efficient model compression to arbitrary sizes with minimal quality loss. (3) Inference-time Scaling: A repeated sampling strategy that trades computation for model capacity, enabling smaller models to match larger model quality at inference time. Through these strategies, SANA-1.5 achieves a text-image alignment score of 0.72 on GenEval, which can be further improved to 0.80 through inference scaling, establishing a new SoTA on GenEval benchmark. These innovations enable efficient model scaling across different compute budgets while maintaining high quality, making high-quality image generation more accessible. Our code and pre-trained models will be released.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc4f5c53-0598-4d4b-8d86-b055f0998461Cited by top-tier papers76
- Flow-GRPO: Training Flow Matching Models via Online RLJie Liu, Gongye Liu, Jiajun Liang, Yangguang Li et al.NeurIPS 2025 · 647 citations
- Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware PermutationShuo Yang, Haocheng Xi, Yilong Zhao, Muyang Li et al.NeurIPS 2025 · 114 citations
- TEMPFLOW-GRPO: WHEN TIMING MATTERS FOR GRPO IN FLOW MODELSXiaoxuan He, Siming Fu, Yuke Zhao, Wanli Li et al.ICLR 2026 · 98 citations
- SANA-Video: Efficient Video Generation with Block Linear Diffusion TransformerJunsong Chen, Yuyang Zhao, Jincheng Yu, Ruihang Chu et al.ICLR 2026 · 96 citations
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at ScaleChunrui Han, Guopeng Li, Jingwei Wu, Quan Sun et al.ICLR 2026 · 58 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- SANA: Efficient High-Resolution Text-to-Image Synthesis with Linear Diffusion TransformersEnze Xie, Junsong Chen, Junyu Chen, Han Cai et al.ICLR 2025
- On the Scalability of Diffusion-based Text-to-Image GenerationHao Li, Yang Zou, Ying Wang, Orchid Majumder et al.CVPR 2024
- Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context ReflectionShufan Li, Konstantinos Kallidromitis, Akash Gokul, Arsh Koneru et al.ICCV 2025 · 3 citations
- EC-DIT: Scaling Diffusion Transformers with Adaptive Expert-Choice RoutingHaotian Sun, Tao Lei, Bowen Zhang, Yanghao Li et al.ICLR 2025
- SANA-Sprint: One-Step Diffusion with Continuous-Time Consistency DistillationJunsong Chen, Shuchen Xue, Yuyang Zhao, Jincheng Yu et al.ICCV 2025 · 6 citations
