Text-to-Vector Generation with Neural Path Representation
Peiying Zhang, Nanxuan Zhao, Jing Liao
Abstract
Vector graphics are widely used in digital art and highly favored by designers due to their scalability and layer-wise properties. However, the process of creating and editing vector graphics requires creativity and design expertise, making it a time-consuming task. Recent advancements in text-to-vector (T2V) generation have aimed to make this process more accessible. However, existing T2V methods directly optimize control points of vector graphics paths, often resulting in intersecting or jagged paths due to the lack of geometry constraints. To overcome these limitations, we propose a novel neural path representation by designing a dual-branch Variational Autoencoder (VAE) that learns the path latent space from both sequence and image modalities. By optimizing the combination of neural paths, we can incorporate geometric constraints while preserving expressivity in generated SVGs. Furthermore, we introduce a two-stage path optimization method to improve the visual and topological quality of generated SVGs. In the first stage, a pre-trained text-to-image diffusion model guides the initial generation of complex vector graphics through the Variational Score Distillation (VSD) process. In the second stage, we refine the graphics using a layer-wise image vectorization strategy to achieve clearer elements and structure. We demonstrate the effectiveness of our method through extensive experiments and showcase various applications. The project page is https://intchous.github.io/T2V-NPR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3cc5eef7-9798-480e-b973-d7e3b985edb0Cited by top-tier papers17
- VAnim: Rendering-Aware Sparse State Modeling for Structure-Preserving Vector AnimationGuotao Liang, Zhangcheng Wang, Chuang Wang, Juncheng Hu et al.ICML 2026 · 11 citations
- LottieGPT: Tokenizing Vector Animation for Autoregressive GenerationJunhao Chen, Kejun Gao, Yuehan Cui, Mingze Sun et al.CVPR 2026 · 10 citations
- SVGen: Interpretable Vector Graphics Generation with Large Language ModelsFeiyu Wang, Zhiyuan Zhao, Yuandong Liu, Da Zhang et al.ACM MM 2025 · 7 citations
- DuetSVG: Unified Multimodal SVG Generation with Internal Visual GuidancePeiying Zhang, Nanxuan Zhao, Matthew Fisher, Yiran Xu et al.CVPR 2026 · 6 citations
- NeuralSVG: An Implicit Representation for Text-to-Vector GenerationSagi Polaczek, Yuval Alaluf, Elad Richardson, Yael Vinker et al.ICCV 2025 · 6 citations
Builds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Style Customization of Text-to-Vector Generation with Image Diffusion PriorsPeiying Zhang, Nanxuan Zhao, Jing LiaoSIGGRAPH 2025
- SVGDreamer: Text Guided SVG Generation with Diffusion ModelXiming Xing, Haitao Zhou, Chuang Wang, Jing Zhang et al.CVPR 2024 · 22 citations
- Im2Vec: Synthesizing Vector Graphics Without Vector SupervisionPradyumna Reddy, Michaël Gharbi, Michal Lukác, Niloy J. MitraCVPR 2021
- DiffSketcher: Text Guided Vector Sketch Synthesis through Latent Diffusion ModelsXiming Xing, Chuang Wang, Haitao Zhou, Jing Zhang et al.NeurIPS 2023 · 101 citations
- LayerTracer: Cognitive-Aligned Layered SVG Synthesis via Diffusion TransformerYiren Song, Danze Chen, Mike Zheng ShouICCV 2025 · 5 citations
