SVGThinker: Instruction-Aligned and Reasoning-Driven Text-to-SVG Generation
Hanqi Chen, Zhongyin Zhao, Ye Chen, Zhujin Liang, Bingbing Ni
Abstract
Scalable Vector Graphics (SVG) is a code structure used to represent visual information, and with the powerful capabilities of large language models, it holds significant research potential. Current text-to-SVG generation methods lack generalization capabilities and struggle with accurately adhering to input generation instructions. In this paper, we propose a novel approach for generating SVG using large language models, named SVGThinker, which incorporates a reasoning process to align the generation of SVG code with the visualization process, while supporting all SVG primitives. Through sequential rendering of SVG primitives, we first use a multimodal model to annotate the SVG, followed by sequential updates corresponding to the incremental additions of primitives. We then employ a supervised training framework based on Chain-of-Thought reasoning, which enhances the model's robustness and reduces the risk of errors or hallucinations. Through comparisons with state-of-the-art baseline models, our experiments show that our model generates more stable, high-quality, and editable SVG code. In contrast to image-based methods, our approach preserves the structural advantages of SVG and supports precise, hierarchical editing. We believe our work opens new directions for SVG generation, with potential applications in design, content creation, and automated SVG-based graphic generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 377a7acf-77b1-44f4-b4e0-16f7a1c84101Cited by top-tier papers2
- VAnim: Rendering-Aware Sparse State Modeling for Structure-Preserving Vector AnimationGuotao Liang, Zhangcheng Wang, Chuang Wang, Juncheng Hu et al.ICML 2026 · 11 citations
- IntroSVG: Learning from Rendering Feedback for Text-to-SVG Generation via an Introspective Generator–Critic FrameworkFeiyu Wang, Jiayuan Yang, Zhiyuan Zhao, Da Zhang et al.CVPR 2026 · 3 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- StarVector: Generating Scalable Vector Graphics Code from Images and TextJuan A. Rodríguez, Abhay Puri, Shubham Agarwal, Issam H. Laradji et al.CVPR 2025
- OmniSVG: A Unified Scalable Vector Graphics Generation ModelYiying Yang, Wei Cheng, Sijin Chen, Xianfang Zeng et al.NeurIPS 2025 · 90 citations
- Rendering-Aware Reinforcement Learning for Vector Graphics GenerationJuan A. Rodríguez, Haotian Zhang, Abhay Puri, Rishav Pramanik et al.NeurIPS 2025 · 42 citations
- Vector Calligrapher: Generating Scalable Vector Graphics via Structured Linguistic SupervisionBo Zhou, Xikang Chen, Yan Gong, Yin ZhangACL 2026
- DuetSVG: Unified Multimodal SVG Generation with Internal Visual GuidancePeiying Zhang, Nanxuan Zhao, Matthew Fisher, Yiran Xu et al.CVPR 2026 · 6 citations
