FontDiffuser: One-Shot Font Generation via Denoising Diffusion with Multi-Scale Content Aggregation and Style Contrastive Learning
Zhenhua Yang, Dezhi Peng, Yuxin Kong, Yuyi Zhang, Cong Yao, Lianwen Jin
摘要
Automatic font generation is an imitation task, which aims to create a font library that mimics the style of reference images while preserving the content from source images. Although existing font generation methods have achieved satisfactory performance, they still struggle with complex characters and large style variations. To address these issues, we propose FontDiffuser, a diffusion-based image-to-image one-shot font generation method, which innovatively models the font imitation task as a noise-to-denoise paradigm. In our method, we introduce a Multi-scale Content Aggregation (MCA) block, which effectively combines global and local content cues across different scales, leading to enhanced preservation of intricate strokes of complex characters. Moreover, to better manage the large variations in style transfer, we propose a Style Contrastive Refinement (SCR) module, which is a novel structure for style representation learning. It utilizes a style extractor to disentangle styles from images, subsequently supervising the diffusion model via a meticulously designed style contrastive loss. Extensive experiments demonstrate FontDiffuser's state-of-the-art performance in generating diverse characters and styles. It consistently excels on complex characters and large style changes compared to previous methods. The code is available at https://github.com/yeungchenwa/FontDiffuser .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- IF-Font: Ideographic Description Sequence-Following Font GenerationXinping Chen, Xiao Ke, Wenzhong GuoNeurIPS 2024 · 被引用 13 次
- Hand1000: Generating Realistic Hands from Text with Only 1, 000 ImagesHaozhuo Zhang, Bin Zhu, Yu Cao, Yanbin HaoAAAI 2025 · 被引用 11 次
- Deciphering Oracle Bone Language with Diffusion ModelsHaisu Guan, Huanxin Yang, Xinyu Wang, Shengwei Han 等ACL 2024 · 被引用 11 次
- Predicting the Original Appearance of Damaged Historical DocumentsZhenhua Yang, Dezhi Peng, Yongxin Shi, Yuyi Zhang 等AAAI 2025 · 被引用 8 次
- FontCraft: Multimodal Font Design Using Interactive Bayesian OptimizationYuki Tatsukawa, I-Chao Shen, Mustafa Doga Dogan, Anran Qi 等CHI 2025 · 被引用 6 次
它引用的顶会 Paper15
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam 等ICML 2022 · 被引用 4,691 次
相关 Paper
- Generate Like Experts: Multi-Stage Font Generation by Incorporating Font Transfer Process into Diffusion ModelsBin Fu, Fanghua Yu, Anran Liu, Zixuan Wang 等CVPR 2024
- Fontanimate: High Quality Few-Shot Font Generation Via Animating Font Transfer ProcessBin Fu, Zixuan Wang, Kainan Yan, Shitian Zhao 等ICCV 2025 · 被引用 1 次
- Decoupling Layout from Glyph in Online Chinese Handwriting GenerationMinsi Ren, Yan-Ming Zhang, Yi ChenICLR 2025
- Multiple Heads are Better than One: Few-shot Font Generation with Multiple Localized ExpertsSong Park, Sanghyuk Chun, Junbum Cha, Bado Lee 等ICCV 2021 · 被引用 96 次
- Look Closer to Supervise Better: One-Shot Font Generation via Component-Based DiscriminatorYuxin Kong, Canjie Luo, Weihong Ma, Qiyuan Zhu 等CVPR 2022 · 被引用 68 次
