FontDiffuser: One-Shot Font Generation via Denoising Diffusion with Multi-Scale Content Aggregation and Style Contrastive Learning
Zhenhua Yang, Dezhi Peng, Yuxin Kong, Yuyi Zhang, Cong Yao, Lianwen Jin
Abstract
Automatic font generation is an imitation task, which aims to create a font library that mimics the style of reference images while preserving the content from source images. Although existing font generation methods have achieved satisfactory performance, they still struggle with complex characters and large style variations. To address these issues, we propose FontDiffuser, a diffusion-based image-to-image one-shot font generation method, which innovatively models the font imitation task as a noise-to-denoise paradigm. In our method, we introduce a Multi-scale Content Aggregation (MCA) block, which effectively combines global and local content cues across different scales, leading to enhanced preservation of intricate strokes of complex characters. Moreover, to better manage the large variations in style transfer, we propose a Style Contrastive Refinement (SCR) module, which is a novel structure for style representation learning. It utilizes a style extractor to disentangle styles from images, subsequently supervising the diffusion model via a meticulously designed style contrastive loss. Extensive experiments demonstrate FontDiffuser's state-of-the-art performance in generating diverse characters and styles. It consistently excels on complex characters and large style changes compared to previous methods. The code is available at https://github.com/yeungchenwa/FontDiffuser .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8e2ae798-74c4-4019-80be-1264bc151a0fCited by top-tier papers20
- IF-Font: Ideographic Description Sequence-Following Font GenerationXinping Chen, Xiao Ke, Wenzhong GuoNeurIPS 2024 · 13 citations
- Hand1000: Generating Realistic Hands from Text with Only 1, 000 ImagesHaozhuo Zhang, Bin Zhu, Yu Cao, Yanbin HaoAAAI 2025 · 11 citations
- Deciphering Oracle Bone Language with Diffusion ModelsHaisu Guan, Huanxin Yang, Xinyu Wang, Shengwei Han et al.ACL 2024 · 11 citations
- Predicting the Original Appearance of Damaged Historical DocumentsZhenhua Yang, Dezhi Peng, Yongxin Shi, Yuyi Zhang et al.AAAI 2025 · 8 citations
- FontCraft: Multimodal Font Design Using Interactive Bayesian OptimizationYuki Tatsukawa, I-Chao Shen, Mustafa Doga Dogan, Anran Qi et al.CHI 2025 · 6 citations
Builds on15
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
Related papers
- Generate Like Experts: Multi-Stage Font Generation by Incorporating Font Transfer Process into Diffusion ModelsBin Fu, Fanghua Yu, Anran Liu, Zixuan Wang et al.CVPR 2024
- Fontanimate: High Quality Few-Shot Font Generation Via Animating Font Transfer ProcessBin Fu, Zixuan Wang, Kainan Yan, Shitian Zhao et al.ICCV 2025 · 1 citation
- Decoupling Layout from Glyph in Online Chinese Handwriting GenerationMinsi Ren, Yan-Ming Zhang, Yi ChenICLR 2025
- Multiple Heads are Better than One: Few-shot Font Generation with Multiple Localized ExpertsSong Park, Sanghyuk Chun, Junbum Cha, Bado Lee et al.ICCV 2021 · 96 citations
- Look Closer to Supervise Better: One-Shot Font Generation via Component-Based DiscriminatorYuxin Kong, Canjie Luo, Weihong Ma, Qiyuan Zhu et al.CVPR 2022 · 68 citations
