FonTS: Text Rendering with Typography and Style Controls
Wenda Shi, Yiren Song, Dengming Zhang, Jiaming Liu, Xingxing Zou
Abstract
Visual text rendering are widespread in various real-world applications, requiring careful font selection and typographic choices. Recent progress in diffusion transformer (DiT)-based text-to-image (T2I) models show promise in automating these processes. However, these methods still encounter challenges like inconsistent fonts, style variation, and limited fine-grained control, particularly at the word-level. This paper proposes a two-stage DiT-based pipeline to address these problems by enhancing controllability over typography and style in text rendering. We introduce typography control fine-tuning (TC-FT), an parameter-efficient fine-tuning method (on key parameters) with enclosing typography control tokens (ETC-tokens), which enables precise word-level application of typographic features. To further address style inconsistency in text rendering, we propose a text-agnostic style control adapter (SCA) that prevents content leakage while enhancing style consistency. To implement TC-FT and SCA effectively, we incorporated HTML-render into the data synthesis pipeline and proposed the first word-level controllable dataset. Through comprehensive experiments, we demonstrate the effectiveness of our approach in achieving superior word-level typographic control, font consistency, and style consistency in text rendering tasks. The datasets and models will be available for academic use.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- RelationAdapter: Learning and Transferring Visual Relation with Diffusion TransformersYan Gong, Yiren Song, Yicheng Li, Chenglin Li et al.NeurIPS 2025 · 30 citations
- EasyText: Controllable Diffusion Transformer for Multilingual Text RenderingRunnan Lu, Yuxuan Zhang, Jiaming Liu, Haofan Wang et al.AAAI 2026 · 20 citations
- InfiniDreamer: Arbitrarily Long Human Motion Generation Via Segment Score DistillationWenjie Zhuo, Fan Ma, Hehe FanICCV 2025 · 6 citations
- EEdit ⚡: Rethinking the Spatial and Temporal Redundancy for Efficient Image EditingZexuan Yan, Yue Ma, Chang Zou, Wenteng Chen et al.ICCV 2025 · 5 citations
- LayerTracer: Cognitive-Aligned Layered SVG Synthesis via Diffusion TransformerYiren Song, Danze Chen, Mike Zheng ShouICCV 2025 · 5 citations
Builds on44
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- ViType: High-Fidelity Visual Text Rendering via Glyph-Aware Multimodal DiffusionLishuai Gao, Jun-Yan He, Yingsen Zeng, Yujie Zhong et al.AAAI 2026
- TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance ControlWeichao Zeng, Yan Shu, Zhenhang Li, Dongbao Yang et al.NeurIPS 2024 · 55 citations
- FreeText: Training-Free Text Rendering via Attention Localization and Spectral Glyph InjectionRuiQiang Zhang, Hengyi Wang, Chang Liu, Guanjie Wang et al.ICML 2026 · 3 citations
- ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion ModelingYuxuan Jiang, Zehua Chen, Zeqian Ju, Yusheng Dai et al.ACL 2026 · 8 citations
- Rethinking Cross-Modal Interaction in Multimodal Diffusion TransformersZhengyao Lv, Tianlin Pan, Chenyang Si, Zhaoxi Chen et al.ICCV 2025 · 3 citations
