Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font Generation
Haonan Cai, Yuxuan Luo, Zhouhui Lian
Abstract
Manual font design is an intricate process that transforms a stylistic visual concept into a coherent glyph set. This challenge persists in automated Few-shot Font Generation (FFG), where models often struggle to preserve both the structural integrity and stylistic fidelity from limited references. While autoregressive (AR) models have demonstrated impressive generative capabilities, their application to FFG is constrained by conventional patch-level tokenization, which neglects global dependencies crucial for coherent font synthesis. Moreover, existing FFG methods remain within the image-to-image paradigm, relying solely on visual references and overlooking the role of language in conveying stylistic intent during font design. To address these limitations, we propose GAR-Font, a novel AR framework for multimodal few-shot font generation. GAR-Font introduces a global-aware tokenizer that effectively captures both local structures and global stylistic patterns, a multimodal style encoder offering flexible style control through a lightweight language-style adapter without requiring intensive multimodal pretraining, and a post-refinement pipeline that further enhances structural fidelity and style coherence. Extensive experiments show that GAR-Font outperforms existing FFG methods, excelling in maintaining global style faithfulness and achieving higher-quality results with textual stylistic guidance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a1eb9ed1-d6a9-4f87-a2d8-7a165c7d0b05Builds on39
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- Vector-quantized Image Modeling with Improved VQGANJiahui Yu, Xin Li, Jing Yu Koh, Han Zhang et al.ICLR 2022 · 753 citations
Related papers
- Multiple Heads are Better than One: Few-shot Font Generation with Multiple Localized ExpertsSong Park, Sanghyuk Chun, Junbum Cha, Bado Lee et al.ICCV 2021 · 96 citations
- VQ-FONT: Few-Shot Font Generation with Structure-Aware Enhancement and QuantizationMingshuai Yao, Yabo Zhang, Xianhui Lin, Xiaoming Li et al.AAAI 2024 · 26 citations
- DA-Font: Few-Shot Font Generation via Dual-Attention Hybrid IntegrationWeiran Chen, Guiqian Zhu, Ying Li, Yi Ji et al.ACM MM 2025 · 2 citations
- Look Closer to Supervise Better: One-Shot Font Generation via Component-Based DiscriminatorYuxin Kong, Canjie Luo, Weihong Ma, Qiyuan Zhu et al.CVPR 2022 · 68 citations
- Few-Shot Font Generation by Learning Fine-Grained Local StylesLicheng Tang, Yiyang Cai, Jiaming Liu, Zhibin Hong et al.CVPR 2022 · 77 citations
