TextCraftor: Your Text Encoder can be Image Quality Controller
Yanyu Li, Xian Liu, Anil Kag, Ju Hu, Yerlan Idelbayev, Dhritiman Sagar, Yanzhi Wang, Sergey Tulyakov, Jian Ren
Abstract
a cartoon of a house on a mountain a cartoon of a boy playing with a tiger an owl standing on a telephone wire a frustrated child world's best brother t-shirt a girl with long curly blonde hair and sunglasses a bowl with a cartoon dinosaur on it a thumbnail image of a gingerbread man a plate with white rice topped by cooked vegetables Figure 1. Example generated images. For each prompt, we show images generated from three different models, which are SDv1.5, TextCraftor, TextCraftor + UNet, listed from left to right. The random seed is fixed for all generation results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6f220e3c-46a6-432b-867e-3860365e5029Cited by top-tier papers12
- ReNO: Enhancing One-step Text-to-Image Models through Reward-based Noise OptimizationLuca Eyring, Shyamgopal Karthik, Karsten Roth, Alexey Dosovitskiy et al.NeurIPS 2024 · 131 citations
- Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion ModelsLuca Eyring, Shyamgopal Karthik, Alexey Dosovitskiy, Nataniel Ruiz et al.NeurIPS 2025 · 36 citations
- G-Refine: A General Quality Refiner for Text-to-Image GenerationChunyi Li, Haoning Wu, Hongkun Hao, Zicheng Zhang et al.ACM MM 2024 · 7 citations
- Enhancing Motion in Text-to-Video Generation with Decomposed Encoding and ConditioningPenghui Ruan, Pichao Wang, Divya Saxena, Jiannong Cao et al.NeurIPS 2024 · 5 citations
- Enhancing Diffusion-based Unrestricted Adversarial Attacks via Adversary Preferences AlignmentKaixun Jiang, Zhaoyu Chen, Haijing Guo, Jinglun Li et al.NeurIPS 2025 · 4 citations
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- StyleDrop: Text-to-Image Synthesis of Any StyleKihyuk Sohn, Lu Jiang, Jarred Barber, Kimin Lee et al.NeurIPS 2023 · 71 citations
- Make It Count: Text-to-Image Generation with an Accurate Number of ObjectsLital Binyamin, Yoad Tewel, Hilit Segev, Eran Hirsch et al.CVPR 2025
- Latent-NeRF for Shape-Guided Generation of 3D Shapes and TexturesGal Metzer, Elad Richardson, Or Patashnik, Raja Giryes et al.CVPR 2023
- Enhancing Compositional Text-to-Image Generation with Reliable Random SeedsShuangqi Li, Hieu Le, Jingyi Xu, Mathieu SalzmannICLR 2025
- ArtAdapter: Text-to-Image Style Transfer using Multi-Level Style Encoder and Explicit AdaptationDar-Yen Chen, Hamish Tennent, Ching-Wen HsuCVPR 2024
