DS-Fusion: Artistic Typography via Discriminated and Stylized Diffusion
Maham Tanveer, Yizhi Wang, Ali Mahdavi-Amiri, Hao Zhang
Abstract
We introduce a novel method to automatically generate an artistic typography by stylizing one or more letter fonts to visually convey the semantics of an input word, while ensuring that the output remains readable. To address an assortment of challenges with our task at hand including conflicting goals (artistic stylization vs. legibility), lack of ground truth, and immense search space, our approach utilizes large language models to bridge texts and visual images for stylization and build an unsupervised generative model with a diffusion model backbone. Specifically, we employ the denoising generator in Latent Diffusion Model (LDM), with the key addition of a CNN-based discriminator to adapt the input style onto the input text. The discriminator uses rasterized images of a given letter/word font as real samples and the output of the denoising generator as fake samples. Our model is coined DS-Fusion for discriminated and stylized diffusion. We showcase the quality and versatility of our method through numerous examples, qualitative and quantitative evaluation, and ablation studies. User studies comparing to strong baselines including CLIPDraw, DALL-E 2, Stable Diffusion, as well as artist-crafted typographies, demonstrate strong performance of DS-Fusion. Code is available at https://ds-fusion.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 32d1aa47-9e8d-4b06-bb6f-b763b01cd4bbCited by top-tier papers10
- Noise-free Score DistillationOren Katzir, Or Patashnik, Daniel Cohen-Or, Dani LischinskiICLR 2024 · 101 citations
- TypeDance: Creating Semantic Typographic Logos from Image through Personalized GenerationShishi Xiao, Liangwei Wang, Xiaojuan Ma, Wei ZengCHI 2024 · 35 citations
- Dynamic Typography: Bringing Text to Life via Video Diffusion PriorZichen Liu, Yihao Meng, Hao Ouyang, Yue Yu et al.ICCV 2025 · 15 citations
- DG3D: Generating High Quality 3D Textured Shapes by Learning to Discriminate Multi-Modal Diffusion-RenderingsQi Zuo, Yafei Song, Jianfang Li, Lin Liu et al.ICCV 2023 · 4 citations
- FonTS: Text Rendering with Typography and Style ControlsWenda Shi, Yiren Song, Dengming Zhang, Jiaming Liu et al.ICCV 2025 · 4 citations
Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Improved Techniques for Training Score-Based Generative ModelsYang Song, Stefano ErmonNeurIPS 2020 · 1,527 citations
Related papers
- Word-As-Image for Semantic TypographyShir Iluz, Yael Vinker, Amir Hertz, Daniel Berio et al.SIGGRAPH 2023 · 67 citations
- ControlStyle: Text-Driven Stylized Image Generation Using Diffusion PriorsJingwen Chen, Yingwei Pan, Ting Yao, Tao MeiACM MM 2023 · 45 citations
- FontDiffuser: One-Shot Font Generation via Denoising Diffusion with Multi-Scale Content Aggregation and Style Contrastive LearningZhenhua Yang, Dezhi Peng, Yuxin Kong, Yuyi Zhang et al.AAAI 2024 · 90 citations
- Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image SynthesisBingda Tang, Boyang Zheng, Sayak Paul, Saining XieCVPR 2025
- DiffDis: Empowering Generative Diffusion Model with Cross-Modal Discrimination CapabilityRunhui Huang, Jianhua Han, Guansong Lu, Xiaodan Liang et al.ICCV 2023 · 10 citations
