XMP-Font: Self-Supervised Cross-Modality Pre-training for Few-Shot Font Generation
Wei Liu, Fangyue Liu, Fei Ding, Qian He, Zili Yi
摘要
Generating a new font library is a very labor-intensive and time-consuming job for glyph-rich scripts. Few-shot font generation is thus required, as it requires only a few glyph references without fine-tuning during test. Existing methods follow the style-content disentanglement paradigm and expect novel fonts to be produced by combining the style codes of the reference glyphs and the content representations of the source. However, these few-shot font generation methods either fail to capture content-independent style representations, or employ localized component-wise style representations, which is insufficient to model many Chinese font styles that involve hyper-component features such as inter-component spacing and “connected-stroke”. To resolve these drawbacks and make the style representations more reliable, we propose a self-supervised cross-modality pre-training strategy and a cross-modality transformer-based encoder that is conditioned jointly on the glyph image and the corresponding stroke labels. The cross-modality encoder is pre-trained in a self-supervised manner to allow effective capture of cross- and intra-modality correlations, which facilitates the content-style disentanglement and modeling style representations of all scales (strokelevel, component-level and character-level). The pretrained encoder is then applied to the downstream font generation task without fine-tuning. Experimental comparisons of our method with state-of-the-art methods demonstrate our method successfully transfers styles of all scales. In addition, it only requires one reference glyph and achieves the lowest rate of bad cases in the few-shot font generation task (28% lower than the second best).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- FontDiffuser: One-Shot Font Generation via Denoising Diffusion with Multi-Scale Content Aggregation and Style Contrastive LearningZhenhua Yang, Dezhi Peng, Yuxin Kong, Yuyi Zhang 等AAAI 2024 · 被引用 90 次
- VQ-FONT: Few-Shot Font Generation with Structure-Aware Enhancement and QuantizationMingshuai Yao, Yabo Zhang, Xianhui Lin, Xiaoming Li 等AAAI 2024 · 被引用 26 次
- Few shot font generation via transferring similarity guided global style and quantization local styleWei Pan, Anna Zhu, Xinyu Zhou, Brian Kenji Iwana 等ICCV 2023 · 被引用 24 次
- IF-Font: Ideographic Description Sequence-Following Font GenerationXinping Chen, Xiao Ke, Wenzhong GuoNeurIPS 2024 · 被引用 13 次
- DeepCalliFont: Few-Shot Chinese Calligraphy Font Synthesis by Integrating Dual-Modality Generative ModelsYitian Liu, Zhouhui LianAAAI 2024 · 被引用 12 次
它引用的顶会 Paper10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- Few-Shot Unsupervised Image-to-Image TranslationMing-Yu Liu, Xun Huang, Arun Mallya, Tero Karras 等ICCV 2019 · 被引用 668 次
- ERNIE-ViL: Knowledge Enhanced Vision-Language Representations through Scene GraphsFei Yu, Jiji Tang, Weichong Yin, Yu Sun 等AAAI 2021 · 被引用 414 次
- Few-shot Font Generation with Localized Style Representations and FactorizationSong Park, Sanghyuk Chun, Junbum Cha, Bado Lee 等AAAI 2021 · 被引用 111 次
相关 Paper
- Few-Shot Font Generation by Learning Fine-Grained Local StylesLicheng Tang, Yiyang Cai, Jiaming Liu, Zhibin Hong 等CVPR 2022 · 被引用 77 次
- Multiple Heads are Better than One: Few-shot Font Generation with Multiple Localized ExpertsSong Park, Sanghyuk Chun, Junbum Cha, Bado Lee 等ICCV 2021 · 被引用 96 次
- Neural Transformation Fields for Arbitrary-Styled Font GenerationBin Fu, Junjun He, Jianjun Wang, Yu QiaoCVPR 2023
- ZiGAN: Fine-grained Chinese Calligraphy Font Generation via a Few-shot Style Transfer ApproachQi Wen, Shuang Li, Bingfeng Han, Yi YuanACM MM 2021 · 被引用 42 次
- DA-Font: Few-Shot Font Generation via Dual-Attention Hybrid IntegrationWeiran Chen, Guiqian Zhu, Ying Li, Yi Ji 等ACM MM 2025 · 被引用 2 次
