Scaling Down Text Encoders of Text-to-Image Diffusion Models
Lifu Wang, Daqing Liu, Xinchen Liu, Xiaodong He
Abstract
Text encoders in diffusion models have rapidly evolved, transitioning from CLIP to T5-XXL. Although this evolution has significantly enhanced the models' ability to understand complex prompts and generate text, it also leads to a substantial increase in the number of parameters. Despite T5 series encoders being trained on the C4 natural language corpus, which includes a significant amount of non-visual data, diffusion models with T5 encoder do not respond to those non-visual prompts, indicating redundancy in representational power. Therefore, it raises an important question: "Do we really need such a large text encoder?" In pursuit of an answer, we employ vision-based knowledge distillation to train a series of T5 encoder models. To fully inherit T5-XXL's capabilities, we constructed our dataset based on three criteria: image quality, semantic understanding, and text-rendering. Our results demonstrate the scaling down pattern that the distilled T5-base model can generate images of comparable quality to those produced by T5-XXL, while being 50 times smaller in size. This reduction in model size significantly lowers the GPU requirements for running state-of-the-art models such as FLUX and SD3, making high-quality text-to-image generation more accessible. Our code is available at https: //lifuwang-66.github.io/ScalingDownTE/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe387ab8-067a-4fbb-878c-fb4341a72f59Cited by top-tier papers4
- Neodragon: Mobile Video Generation Using Diffusion TransformerAnimesh Karnewar, Denis Korzhenkov, Ioannis Lelekas, Noor Fathima et al.ICLR 2026 · 10 citations
- Beyond the Seen: Bounded Distribution Estimation for Open-Vocabulary LearningXiaomeng Fan, Yuchuan Mao, Zhi Gao, Yuwei Wu et al.NeurIPS 2025 · 1 citation
- Curriculum Group Policy Optimization: Adaptive Sampling for Unleashing the Potential of Text-to-Image GenerationBaoteng Li, Xianghao Zang, Xinran Wang, Xiangyu Na et al.CVPR 2026
- NanoFLUX: Distillation-Driven Compression of Large Text-to-Image Generation Models for Mobile DevicesRuchika Chavhan, Malcolm Chadwick, Alberto Gil Couto Pimentel Ramos, Luca Morreale et al.ICML 2026
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- A Comprehensive Study of Decoder-Only LLMs for Text-to-Image GenerationAndrew Z. Wang, Songwei Ge, Tero Karras, Ming-Yu Liu et al.CVPR 2025
- KOALA: Empirical Lessons Toward Memory-Efficient and Fast Diffusion Models for Text-to-Image SynthesisYoungwan Lee, Kwanyong Park, Yoorhim Cho, Yong-Ju Lee et al.NeurIPS 2024 · 22 citations
- SnapFusion: Text-to-Image Diffusion Model on Mobile Devices within Two SecondsYanyu Li, Huan Wang, Qing Jin, Ju Hu et al.NeurIPS 2023 · 300 citations
- On the Scalability of Diffusion-based Text-to-Image GenerationHao Li, Yang Zou, Ying Wang, Orchid Majumder et al.CVPR 2024
- Diffusion Lens: Interpreting Text Encoders in Text-to-Image PipelinesMichael Toker, Hadas Orgad, Mor Ventura, Dana Arad et al.ACL 2024
