InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models
Haomin Wang, Jinhui Yin, Qi Wei, Wenguang Zeng, Lixin Gu, Shenglong Ye, Zhangwei Gao, Yaohui Wang, Yanting Zhang, Yuanqi Li, Yanwen Guo, Wenhai Wang
摘要
Published as a conference paper at ICLR 2026 tokens for tags, attributes, and coordinates to reduce sequence length while retaining geometric and hierarchical structure. These tokens are initialized with a subword-based strategy that anchors them in the pretrained embedding space, stabilizing early training and accelerating convergence. Training adopts a two-stage strategy that progresses from short static SVGs to longer illustrations and complex animations. Through extensive experiments, we demonstrate that unified modeling can effectively improve performance across understanding, editing, and generation tasks. Comprehensive evaluations further show that our InternSVG surpasses both open-source and proprietary models on SArena and previous benchmarks. For example, on the SArena-Icon benchmark, InternSVG surpasses Claude-Sonnet-4, the strongest proprietary baseline on SVG tasks, by about 11% higher acc in understanding tasks, 34% higher PSNR in editing tasks, 56% lower FID in Text-to-SVG tasks, and 22% higher SSIM in Image-to-SVG tasks. In summary, our contributions are below: (1) We construct SAgoge, the largest and most comprehensive multimodal SVG dataset to date, encompassing static graphics and animations with over 16 million training samples. To enable rigorous and comparable evaluation, we further establish SArena, a companion benchmark that standardizes tasks and metrics across SVG understanding, editing, and generation. (2) We propose InternSVG, a unified MLLM for SVG understanding, editing, and generation. It introduces SVG-specific tokenization with subword-initialized special tokens and adopts a two-stage training strategy to support effective cross-task generalization. (3) We conduct extensive experiments to demonstrate the benefits of unified modeling. The results on SArena and prior benchmarks show that our InternSVG outperforms traditional approaches as well as general-purpose open-source and proprietary models. RELATED WORKS 2.1 SVG DATASETS AND BENCHMARKS Most existing SVG datasets and benchmarks are limited in task coverage or data type and remain too small for effective model training, leading to fragmented evaluations and limited insights into generalization across tasks and complexity. SGP-Bench (Qiu et al., 2024) evaluates semantic comprehension and consistency in symbolic graphics programs. SVGEditBench (Nishina & Matsui, 2024) and its extension V2 (Nishina & Matsui, 2025) focus narrowly on instruction-based SVG editing measured by low-level syntactic metrics. On the generative side, SVG-Stack (Rodriguez et al., 2025), SVGX (Xing et al., 2025), MMSVG (Yang et al., 2025b), and ColorSVG-100K (Chen & Pan, 2025) address Text-to-SVG and Image-to-SVG generation, while VGBench Zou et al. (2024) and UniSVG (Li et al., 2025) jointly evaluate understanding and generation. DeepSVG (Carlier et al., 2020) introduces a dataset of 100K SVG icons and explores generation, interpolation, and latentspace animation of static and limited animated graphics, but lacks rich editing instructions and image-conditioned generation. SVGenius (Chen et al., 2025) introduces a comprehensive benchmark covering understanding, editing, and generation with systematic complexity levels and multidimensional metrics, but it includes only about 2,400 queries, making it sufficient for evaluation yet inadequate for training. In contrast, our SAgoge is substantially larger and more diverse, encompassing both static graphics and SVG animations. It unifies SVG understanding, editing, and generation, and with approximately 16M task samples, its scale and diversity enable robust model training and comprehensive evaluation across the full spectrum of SVG tasks, which effectively address the coverage and scalability limitations of prior datasets. SVG MODELING METHODS Early research on SVG modeling treated vector graphics as sequences of geometric primitives and relied on specialized generative architectures trained on limited domains (
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- DuetSVG: Unified Multimodal SVG Generation with Internal Visual GuidancePeiying Zhang, Nanxuan Zhao, Matthew Fisher, Yiran Xu 等CVPR 2026 · 被引用 6 次
- Vector Prism: Animating Vector Graphics by Stratifying Semantic StructureJooyeol Yun, Jaegul ChooCVPR 2026 · 被引用 2 次
- VectorArk: Learning Practical Image Vectorization with Rounded Polygon RepresentationTarun Gehlaut, Difan Liu, Charu Bansal, Krutik Malani 等CVPR 2026
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras 等EMNLP 2021 · 被引用 937 次
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li 等ICLR 2024 · 被引用 467 次
- CLIPDraw: Exploring Text-to-Drawing Synthesis through Language-Image EncodersKevin Frans, Lisa B. Soros, Olaf WitkowskiNeurIPS 2022 · 被引用 311 次
相关 Paper
- OmniSVG: A Unified Scalable Vector Graphics Generation ModelYiying Yang, Wei Cheng, Sijin Chen, Xianfang Zeng 等NeurIPS 2025 · 被引用 90 次
- StarVector: Generating Scalable Vector Graphics Code from Images and TextJuan A. Rodríguez, Abhay Puri, Shubham Agarwal, Issam H. Laradji 等CVPR 2025
- DeepSVG: A Hierarchical Generative Network for Vector Graphics AnimationAlexandre Carlier, Martin Danelljan, Alexandre Alahi, Radu TimofteNeurIPS 2020 · 被引用 247 次
- SVGThinker: Instruction-Aligned and Reasoning-Driven Text-to-SVG GenerationHanqi Chen, Zhongyin Zhao, Ye Chen, Zhujin Liang 等ACM MM 2025 · 被引用 3 次
- SVGen: Interpretable Vector Graphics Generation with Large Language ModelsFeiyu Wang, Zhiyuan Zhao, Yuandong Liu, Da Zhang 等ACM MM 2025 · 被引用 7 次
