Learning to Generate Stylized Handwritten Text via a Unified Representation of Style, Content, and Noise
Honglie Wang, Yan-Ming Zhang, Wangzi Yao, Fei Yin, Cheng-Lin Liu
摘要
Handwritten Text Generation (HTG) seeks to synthesize realistic and personalized handwriting by modeling stylistic and structural traits. While recent diffusion-based approaches have advanced generation fidelity, they typically rely on auxiliary style or content encoders with handcrafted objectives, leading to complex training pipelines and limited interaction across factors. In this work, we present InkSpire, a diffusion transformer based model that unifies style, content, and noise within a shared latent space. By eliminating explicit encoders, InkSpire streamlines optimization while enabling richer feature interaction and stronger in-context generation. To further enhance flexibility, we introduce a multi-line masked infilling strategy that allows training directly on raw text-line images, together with a revised positional encoding that supports arbitrary-length multi-line synthesis and fine-grained character editing. Moreover, InkSpire is trained on a bilingual Chinese–English corpus, enabling a single model to handle both Chinese and English handwriting generation with high fidelity and stylistic diversity, thereby overcoming the need for language-specific systems. Extensive experiments on IAM and ICDAR2013 demonstrate that InkSpire achieves superior structural accuracy and stylistic diversity compared to prior state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- Look Closer to Supervise Better: One-Shot Font Generation via Component-Based DiscriminatorYuxin Kong, Canjie Luo, Weihong Ma, Qiyuan Zhu 等CVPR 2022 · 被引用 68 次
- XMP-Font: Self-Supervised Cross-Modality Pre-training for Few-Shot Font GenerationWei Liu, Fangyue Liu, Fei Ding, Qian He 等CVPR 2022 · 被引用 64 次
- Handwriting TransformersAnkan Kumar Bhunia, Salman H. Khan, Hisham Cholakkal, Rao Muhammad Anwer 等ICCV 2021 · 被引用 64 次
- Emu Edit: Precise Image Editing via Recognition and Generation TasksShelly Sheynin, Adam Polyak, Uriel Singer, Yuval Kirstain 等CVPR 2024 · 被引用 62 次
- DiffUTE: Universal Text Editing Diffusion ModelHaoxing Chen, Zhuoer Xu, Zhangxuan Gu, Jun Lan 等NeurIPS 2023 · 被引用 61 次
相关 Paper
- DiffInk: Glyph- and Style-Aware Latent Diffusion Transformer for Text to Online Handwriting GenerationWei Pan, Huiguo He, Hiuyi Cheng, Yilin Shi 等ICLR 2026 · 被引用 2 次
- Beyond Isolated Words: Diffusion Brush for Handwritten Text-Line GenerationGang Dai, Yifan Zhang, Yutao Qin, Qiangya Guo 等ICCV 2025 · 被引用 5 次
- Decoupling Layout from Glyph in Online Chinese Handwriting GenerationMinsi Ren, Yan-Ming Zhang, Yi ChenICLR 2025
- Disentangling Writer and Character Styles for Handwriting GenerationGang Dai, Yifan Zhang, Qingfeng Wang, Qing Du 等CVPR 2023
- FreeText: Training-Free Text Rendering via Attention Localization and Spectral Glyph InjectionRuiQiang Zhang, Hengyi Wang, Chang Liu, Guanjie Wang 等ICML 2026 · 被引用 3 次
