SVGThinker: Instruction-Aligned and Reasoning-Driven Text-to-SVG Generation
Hanqi Chen, Zhongyin Zhao, Ye Chen, Zhujin Liang, Bingbing Ni
摘要
Scalable Vector Graphics (SVG) is a code structure used to represent visual information, and with the powerful capabilities of large language models, it holds significant research potential. Current text-to-SVG generation methods lack generalization capabilities and struggle with accurately adhering to input generation instructions. In this paper, we propose a novel approach for generating SVG using large language models, named SVGThinker, which incorporates a reasoning process to align the generation of SVG code with the visualization process, while supporting all SVG primitives. Through sequential rendering of SVG primitives, we first use a multimodal model to annotate the SVG, followed by sequential updates corresponding to the incremental additions of primitives. We then employ a supervised training framework based on Chain-of-Thought reasoning, which enhances the model's robustness and reduces the risk of errors or hallucinations. Through comparisons with state-of-the-art baseline models, our experiments show that our model generates more stable, high-quality, and editable SVG code. In contrast to image-based methods, our approach preserves the structural advantages of SVG and supports precise, hierarchical editing. We believe our work opens new directions for SVG generation, with potential applications in design, content creation, and automated SVG-based graphic generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- VAnim: Rendering-Aware Sparse State Modeling for Structure-Preserving Vector AnimationGuotao Liang, Zhangcheng Wang, Chuang Wang, Juncheng Hu 等ICML 2026 · 被引用 11 次
- IntroSVG: Learning from Rendering Feedback for Text-to-SVG Generation via an Introspective Generator–Critic FrameworkFeiyu Wang, Jiayuan Yang, Zhiyuan Zhao, Da Zhang 等CVPR 2026 · 被引用 3 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- StarVector: Generating Scalable Vector Graphics Code from Images and TextJuan A. Rodríguez, Abhay Puri, Shubham Agarwal, Issam H. Laradji 等CVPR 2025
- OmniSVG: A Unified Scalable Vector Graphics Generation ModelYiying Yang, Wei Cheng, Sijin Chen, Xianfang Zeng 等NeurIPS 2025 · 被引用 90 次
- Rendering-Aware Reinforcement Learning for Vector Graphics GenerationJuan A. Rodríguez, Haotian Zhang, Abhay Puri, Rishav Pramanik 等NeurIPS 2025 · 被引用 42 次
- Vector Calligrapher: Generating Scalable Vector Graphics via Structured Linguistic SupervisionBo Zhou, Xikang Chen, Yan Gong, Yin ZhangACL 2026
- DuetSVG: Unified Multimodal SVG Generation with Internal Visual GuidancePeiying Zhang, Nanxuan Zhao, Matthew Fisher, Yiran Xu 等CVPR 2026 · 被引用 6 次
