NanoFLUX: Distillation-Driven Compression of Large Text-to-Image Generation Models for Mobile Devices
Ruchika Chavhan, Malcolm Chadwick, Alberto Gil Couto Pimentel Ramos, Luca Morreale, Mehdi Noroozi, Abhinav Mehrotra
Abstract
While large-scale text-to-image diffusion models continue to improve in visual quality, their increasing scale has widened the gap between state-of-the-art models and on-device solutions. To address this gap, we introduce NanoFLUX, a 2.4B text-to-image flow-matching model distilled from 17B FLUX.1-Schnell using a progressive compression pipeline designed to preserve generation quality. Our contributions include: (1) A model compression strategy driven by pruning redundant components in the diffusion transformer, reducing its size from 12B to 2B; (2) A ResNet-based token downsampling mechanism that reduces latency by allowing intermediate blocks to operate on lower-resolution tokens while preserving high-resolution processing elsewhere; (3) A novel text encoder distillation approach that leverages visual signals from early layers of the denoiser during sampling. Empirically, NanoFLUX generates images in approximately 2.5 seconds on mobile devices, demonstrating the feasibility of high-quality on-device text-to-image generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9a46a784-9f72-4ca2-bf55-931f981ce05eBuilds on17
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- Towards Faster and Stabilized GAN Training for High-fidelity Few-shot Image SynthesisBingchen Liu, Yizhe Zhu, Kunpeng Song, Ahmed ElgammalICLR 2021 · 307 citations
- SnapFusion: Text-to-Image Diffusion Model on Mobile Devices within Two SecondsYanyu Li, Huan Wang, Qing Jin, Ju Hu et al.NeurIPS 2023 · 300 citations
- DiTFastAttn: Attention Compression for Diffusion Transformer ModelsZhihang Yuan, Hanling Zhang, Lu Pu, Xuefei Ning et al.NeurIPS 2024 · 134 citations
Related papers
- Scaling Down Text Encoders of Text-to-Image Diffusion ModelsLifu Wang, Daqing Liu, Xinchen Liu, Xiaodong HeCVPR 2025
- SANA-Sprint: One-Step Diffusion with Continuous-Time Consistency DistillationJunsong Chen, Shuchen Xue, Yuyang Zhao, Jincheng Yu et al.ICCV 2025 · 6 citations
- NanoSD: Edge Efficient Foundation Model for Real Time Image RestorationSubhajit Sanyal, Srinivas Soumitri Miriyala, Akshay Janardan Bankar, Manjunath Arveti et al.CVPR 2026 · 3 citations
- Multistep Distillation of Diffusion Models via Moment MatchingTim Salimans, Thomas Mensink, Jonathan Heek, Emiel HoogeboomNeurIPS 2024 · 93 citations
- From Sketch to Fresco: Efficient Diffusion Transformer with Progressive ResolutionShikang Zheng, Guantao Chen, Landis He, Jiacheng Liu et al.CVPR 2026 · 6 citations
