Generate Like Experts: Multi-Stage Font Generation by Incorporating Font Transfer Process into Diffusion Models
Bin Fu, Fanghua Yu, Anran Liu, Zixuan Wang, Jie Wen, Junjun He, Yu Qiao
摘要
Few-shot font generation (FFG) produces stylized font images with a limited number of reference samples, which can significantly reduce labor costs in manual font designs. Most existing FFG methods follow the style-content disentanglement paradigm and employ the Generative Adversarial Network (GAN) to generate target fonts by combining the decoupled content and style representations. The complicated structure and detailed style are simultaneously generated in those methods, which may be the sub-optimal solutions for FFG task. Inspired by most manual font design processes of expert designers, in this paper, we model font generation as a multi-stage generative process. Specifically, as the injected noise and the data distribution in diffusion models can be well-separated into different sub-spaces, we are able to incorporate the font transfer process into these models. Based on this observation, we generalize diffusion methods to model font generative process by separating the reverse diffusion process into three stages with different functions: The structure construction stage first generates the structure information for the target character based on the source image, and the font transfer stage subsequently transforms the source font to the target font. Finally, the font refinement stage enhances the appearances and local details of the target font images. Based on the above multistage generative process, we construct our font generation framework, named MSD-Font, with a dual-network approach to generate font images. The superior performance demonstrates the effectiveness of our model. The code is available at: https://github.com/fubinfb/MSD-Font .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Unbiased Missing-Modality Multimodal LearningRuiting Dai, Chenxi Li, Yandong Yan, Lisi Mo 等ICCV 2025 · 被引用 8 次
- FontCraft: Multimodal Font Design Using Interactive Bayesian OptimizationYuki Tatsukawa, I-Chao Shen, Mustafa Doga Dogan, Anran Qi 等CHI 2025 · 被引用 6 次
- Fontanimate: High Quality Few-Shot Font Generation Via Animating Font Transfer ProcessBin Fu, Zixuan Wang, Kainan Yan, Shitian Zhao 等ICCV 2025 · 被引用 1 次
- Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font GenerationHaonan Cai, Yuxuan Luo, Zhouhui LianCVPR 2026 · 被引用 1 次
- Rethinking Glyph Spatial Information in Font GenerationPeng Su, Xi YangCVPR 2026
它引用的顶会 Paper29
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 被引用 5,234 次
相关 Paper
- Neural Transformation Fields for Arbitrary-Styled Font GenerationBin Fu, Junjun He, Jianjun Wang, Yu QiaoCVPR 2023
- Multiple Heads are Better than One: Few-shot Font Generation with Multiple Localized ExpertsSong Park, Sanghyuk Chun, Junbum Cha, Bado Lee 等ICCV 2021 · 被引用 96 次
- Few-Shot Font Generation by Learning Fine-Grained Local StylesLicheng Tang, Yiyang Cai, Jiaming Liu, Zhibin Hong 等CVPR 2022 · 被引用 77 次
- DA-Font: Few-Shot Font Generation via Dual-Attention Hybrid IntegrationWeiran Chen, Guiqian Zhu, Ying Li, Yi Ji 等ACM MM 2025 · 被引用 2 次
- QT-Font: High-efficiency Font Synthesis via Quadtree-based Diffusion ModelsYitian Liu, Zhouhui LianSIGGRAPH 2024 · 被引用 4 次
