VecGlypher: Unified Vector Glyph Generation with Language Models
Xiaoke Huang, Bhavul Gauri, Kam Woh Ng, Tony Ng, Mengmeng Xu, Zhiheng Liu, Weiming Ren, Zhaochong An, Zijian Zhou, Haonan Qiu, Yuyin Zhou, Sen He
Abstract
Vector glyphs are the atomic units of digital typography, yet most learning-based pipelines still depend on carefully curated exemplar sheets and raster-to-vector postprocessing, which limits accessibility and editability. We introduce VecGlypher, a single multimodal language model that generates high-fidelity vector glyphs directly from text descriptions or image exemplars. Given a style prompt, optional reference glyph images, and a target character, VecGlypher autoregressively emits SVG path tokens, avoiding raster intermediates and producing editable, watertight outlines in one pass. A typography-aware data and training recipe makes this possible: (i) a large-scale continuation stage on 39K noisy Envato fonts to master SVG syntax and long-horizon geometry, followed by (ii) post-training on 2.5K expert-annotated Google Fonts with descriptive tags and exemplars to align language and imagery with geometry; preprocessing normalizes coordinate frames, canonicalizes paths, de-duplicates families, and quantizes coordinates for stable long-sequence decoding. On cross-family OOD evaluation, VecGlypher substantially outperforms both general-purpose LLMs and specialized vector-font baselines for text-only generation, while image-referenced generation reaches a state-of-the-art performance, with marked gains over DeepVecFont-v2 and DualVector. Ablations show that model scale and the two-stage recipe are critical and that absolute-coordinate serialization yields the best geometry. VecGlypher lowers the barrier to font creation by letting users design with words or exemplars, and provides a scalable foundation for future multimodal design tools.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- A Learned Representation for Scalable Vector GraphicsRaphael Gontijo Lopes, David Ha, Douglas Eck, Jonathon ShlensICCV 2019 · 153 citations
Related papers
- Vector Calligrapher: Generating Scalable Vector Graphics via Structured Linguistic SupervisionBo Zhou, Xikang Chen, Yan Gong, Yin ZhangACL 2026
- VecDesigner: Exploring Visual Guidance and Structural Consistency for Semantic TypographyLiu Yu, Xingjiao Wu, Ziang Liu, Jiabao Zhao et al.ICML 2026
- VecFontSDF: Learning to Reconstruct and Synthesize High-Quality Vector Fonts via Signed Distance FunctionsZeqing Xia, Bojun Xiong, Zhouhui LianCVPR 2023
- DeepVecFont-v2: Exploiting Transformers to Synthesize Vector Fonts with Higher QualityYuqing Wang, Yizhi Wang, Longhui Yu, Yuesheng Zhu et al.CVPR 2023
- DuetSVG: Unified Multimodal SVG Generation with Internal Visual GuidancePeiying Zhang, Nanxuan Zhao, Matthew Fisher, Yiran Xu et al.CVPR 2026 · 6 citations
