Learning to Generate Stylized Handwritten Text via a Unified Representation of Style, Content, and Noise
Honglie Wang, Yan-Ming Zhang, Wangzi Yao, Fei Yin, Cheng-Lin Liu
Abstract
Handwritten Text Generation (HTG) seeks to synthesize realistic and personalized handwriting by modeling stylistic and structural traits. While recent diffusion-based approaches have advanced generation fidelity, they typically rely on auxiliary style or content encoders with handcrafted objectives, leading to complex training pipelines and limited interaction across factors. In this work, we present InkSpire, a diffusion transformer based model that unifies style, content, and noise within a shared latent space. By eliminating explicit encoders, InkSpire streamlines optimization while enabling richer feature interaction and stronger in-context generation. To further enhance flexibility, we introduce a multi-line masked infilling strategy that allows training directly on raw text-line images, together with a revised positional encoding that supports arbitrary-length multi-line synthesis and fine-grained character editing. Moreover, InkSpire is trained on a bilingual Chinese–English corpus, enabling a single model to handle both Chinese and English handwriting generation with high fidelity and stylistic diversity, thereby overcoming the need for language-specific systems. Extensive experiments on IAM and ICDAR2013 demonstrate that InkSpire achieves superior structural accuracy and stylistic diversity compared to prior state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 67d821bd-4812-4ee9-be83-c8bfd3edc329Cited by top-tier papers1
Ask how each one uses itBuilds on14
- Look Closer to Supervise Better: One-Shot Font Generation via Component-Based DiscriminatorYuxin Kong, Canjie Luo, Weihong Ma, Qiyuan Zhu et al.CVPR 2022 · 68 citations
- XMP-Font: Self-Supervised Cross-Modality Pre-training for Few-Shot Font GenerationWei Liu, Fangyue Liu, Fei Ding, Qian He et al.CVPR 2022 · 64 citations
- Handwriting TransformersAnkan Kumar Bhunia, Salman H. Khan, Hisham Cholakkal, Rao Muhammad Anwer et al.ICCV 2021 · 64 citations
- Emu Edit: Precise Image Editing via Recognition and Generation TasksShelly Sheynin, Adam Polyak, Uriel Singer, Yuval Kirstain et al.CVPR 2024 · 62 citations
- DiffUTE: Universal Text Editing Diffusion ModelHaoxing Chen, Zhuoer Xu, Zhangxuan Gu, Jun Lan et al.NeurIPS 2023 · 61 citations
Related papers
- DiffInk: Glyph- and Style-Aware Latent Diffusion Transformer for Text to Online Handwriting GenerationWei Pan, Huiguo He, Hiuyi Cheng, Yilin Shi et al.ICLR 2026 · 2 citations
- Beyond Isolated Words: Diffusion Brush for Handwritten Text-Line GenerationGang Dai, Yifan Zhang, Yutao Qin, Qiangya Guo et al.ICCV 2025 · 5 citations
- Decoupling Layout from Glyph in Online Chinese Handwriting GenerationMinsi Ren, Yan-Ming Zhang, Yi ChenICLR 2025
- Disentangling Writer and Character Styles for Handwriting GenerationGang Dai, Yifan Zhang, Qingfeng Wang, Qing Du et al.CVPR 2023
- FreeText: Training-Free Text Rendering via Attention Localization and Spectral Glyph InjectionRuiQiang Zhang, Hengyi Wang, Chang Liu, Guanjie Wang et al.ICML 2026 · 3 citations
