XMP-Font: Self-Supervised Cross-Modality Pre-training for Few-Shot Font Generation
Wei Liu, Fangyue Liu, Fei Ding, Qian He, Zili Yi
Abstract
Generating a new font library is a very labor-intensive and time-consuming job for glyph-rich scripts. Few-shot font generation is thus required, as it requires only a few glyph references without fine-tuning during test. Existing methods follow the style-content disentanglement paradigm and expect novel fonts to be produced by combining the style codes of the reference glyphs and the content representations of the source. However, these few-shot font generation methods either fail to capture content-independent style representations, or employ localized component-wise style representations, which is insufficient to model many Chinese font styles that involve hyper-component features such as inter-component spacing and “connected-stroke”. To resolve these drawbacks and make the style representations more reliable, we propose a self-supervised cross-modality pre-training strategy and a cross-modality transformer-based encoder that is conditioned jointly on the glyph image and the corresponding stroke labels. The cross-modality encoder is pre-trained in a self-supervised manner to allow effective capture of cross- and intra-modality correlations, which facilitates the content-style disentanglement and modeling style representations of all scales (strokelevel, component-level and character-level). The pretrained encoder is then applied to the downstream font generation task without fine-tuning. Experimental comparisons of our method with state-of-the-art methods demonstrate our method successfully transfers styles of all scales. In addition, it only requires one reference glyph and achieves the lowest rate of bad cases in the few-shot font generation task (28% lower than the second best).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3637f911-141e-4ac1-87ab-ed8018e71dc9Cited by top-tier papers17
- FontDiffuser: One-Shot Font Generation via Denoising Diffusion with Multi-Scale Content Aggregation and Style Contrastive LearningZhenhua Yang, Dezhi Peng, Yuxin Kong, Yuyi Zhang et al.AAAI 2024 · 90 citations
- VQ-FONT: Few-Shot Font Generation with Structure-Aware Enhancement and QuantizationMingshuai Yao, Yabo Zhang, Xianhui Lin, Xiaoming Li et al.AAAI 2024 · 26 citations
- Few shot font generation via transferring similarity guided global style and quantization local styleWei Pan, Anna Zhu, Xinyu Zhou, Brian Kenji Iwana et al.ICCV 2023 · 24 citations
- IF-Font: Ideographic Description Sequence-Following Font GenerationXinping Chen, Xiao Ke, Wenzhong GuoNeurIPS 2024 · 13 citations
- DeepCalliFont: Few-Shot Chinese Calligraphy Font Synthesis by Integrating Dual-Modality Generative ModelsYitian Liu, Zhouhui LianAAAI 2024 · 12 citations
Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- Few-Shot Unsupervised Image-to-Image TranslationMing-Yu Liu, Xun Huang, Arun Mallya, Tero Karras et al.ICCV 2019 · 668 citations
- ERNIE-ViL: Knowledge Enhanced Vision-Language Representations through Scene GraphsFei Yu, Jiji Tang, Weichong Yin, Yu Sun et al.AAAI 2021 · 414 citations
- Few-shot Font Generation with Localized Style Representations and FactorizationSong Park, Sanghyuk Chun, Junbum Cha, Bado Lee et al.AAAI 2021 · 111 citations
Related papers
- Few-Shot Font Generation by Learning Fine-Grained Local StylesLicheng Tang, Yiyang Cai, Jiaming Liu, Zhibin Hong et al.CVPR 2022 · 77 citations
- Multiple Heads are Better than One: Few-shot Font Generation with Multiple Localized ExpertsSong Park, Sanghyuk Chun, Junbum Cha, Bado Lee et al.ICCV 2021 · 96 citations
- Neural Transformation Fields for Arbitrary-Styled Font GenerationBin Fu, Junjun He, Jianjun Wang, Yu QiaoCVPR 2023
- ZiGAN: Fine-grained Chinese Calligraphy Font Generation via a Few-shot Style Transfer ApproachQi Wen, Shuang Li, Bingfeng Han, Yi YuanACM MM 2021 · 42 citations
- DA-Font: Few-Shot Font Generation via Dual-Attention Hybrid IntegrationWeiran Chen, Guiqian Zhu, Ying Li, Yi Ji et al.ACM MM 2025 · 2 citations
