DuetSVG: Unified Multimodal SVG Generation with Internal Visual Guidance
Peiying Zhang, Nanxuan Zhao, Matthew Fisher, Yiran Xu, Jing Liao, Difan Liu
Abstract
Recent vision–language model (VLM)–based approaches have achieved impressive results on SVG generation. However, because they generate only text and lack visual signals during decoding, they often struggle with complex semantics and fail to produce visually appealing or geometrically coherent SVGs. We introduce DuetSVG, a unified multimodal model that jointly generates image tokens and corresponding SVG tokens in an end-to-end manner. DuetSVG is trained on both image and SVG datasets. At inference, we apply a novel test-time scaling strategy that leverages the model’s native visual predictions as guidance to improve SVG decoding quality. Extensive experiments show that our method outperforms existing methods, producing visually faithful, semantically aligned, and syntactically clean SVGs across a wide range of applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a96ee37c-0b5f-4dcc-bc52-756d314a5f5fCited by top-tier papers1
Ask how each one uses itBuilds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score DistillationZhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao et al.NeurIPS 2023 · 1,498 citations
- DreamFusion: Text-to-3D using 2D DiffusionBen Poole, Ajay Jain, Jonathan T. Barron, Ben MildenhallICLR 2023 · 463 citations
Related papers
- OmniSVG: A Unified Scalable Vector Graphics Generation ModelYiying Yang, Wei Cheng, Sijin Chen, Xianfang Zeng et al.NeurIPS 2025 · 90 citations
- Vector Calligrapher: Generating Scalable Vector Graphics via Structured Linguistic SupervisionBo Zhou, Xikang Chen, Yan Gong, Yin ZhangACL 2026
- SVGThinker: Instruction-Aligned and Reasoning-Driven Text-to-SVG GenerationHanqi Chen, Zhongyin Zhao, Ye Chen, Zhujin Liang et al.ACM MM 2025 · 3 citations
- Rendering-Aware Reinforcement Learning for Vector Graphics GenerationJuan A. Rodríguez, Haotian Zhang, Abhay Puri, Rishav Pramanik et al.NeurIPS 2025 · 42 citations
- Vector Prism: Animating Vector Graphics by Stratifying Semantic StructureJooyeol Yun, Jaegul ChooCVPR 2026 · 2 citations
