Rendering-Aware Reinforcement Learning for Vector Graphics Generation
Juan A. Rodríguez, Haotian Zhang, Abhay Puri, Rishav Pramanik, Aarash Feizi, Pascal Wichmann, Arnab Kumar Mondal, Mohammad Reza Samsami, Rabiul Awal, Perouz Taslakian, Spandana Gella, Sai Rajeswar Mudumba
摘要
Scalable Vector Graphics (SVG) offer a powerful format for representing visual designs as interpretable code. Recent advances in vision-language models (VLMs) have enabled high-quality SVG generation by framing the problem as a code generation task and leveraging large-scale pretraining. VLMs are particularly suitable for this task as they capture both global semantics and fine-grained visual patterns, while transferring knowledge across vision, natural language, and code domains. However, existing VLM approaches often struggle to produce faithful and efficient SVGs because they never observe the rendered images during training. Although differentiable rendering for autoregressive SVG code generation remains unavailable, rendered outputs can still be compared to original inputs, enabling evaluative feedback suitable for reinforcement learning (RL). We introduce RLRF (Reinforcement Learning from Rendering Feedback), an RL method that enhances SVG generation in autoregressive VLMs by leveraging feedback from rendered SVG outputs. Given an input image, the model generates SVG roll-outs that are rendered and compared to the original image to compute a reward. This visual fidelity feedback guides the model toward producing more accurate, efficient, and semantically coherent SVGs. RLRF significantly outperforms supervised fine-tuning, addressing common failure modes and enabling precise, high-quality SVG generation with strong structural understanding and generalization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code GenerationLei Chen, Xuanle Zhao, Zhixiong Zeng, Jing Huang 等ICLR 2026 · 被引用 16 次
- VAnim: Rendering-Aware Sparse State Modeling for Structure-Preserving Vector AnimationGuotao Liang, Zhangcheng Wang, Chuang Wang, Juncheng Hu 等ICML 2026 · 被引用 11 次
- DuetSVG: Unified Multimodal SVG Generation with Internal Visual GuidancePeiying Zhang, Nanxuan Zhao, Matthew Fisher, Yiran Xu 等CVPR 2026 · 被引用 6 次
- TikZilla: Scaling Text-to-TikZ with High-Quality Data and Reinforcement LearningChristian Greisinger, Steffen EgerICLR 2026 · 被引用 5 次
- IntroSVG: Learning from Rendering Feedback for Text-to-SVG Generation via an Introspective Generator–Critic FrameworkFeiyu Wang, Jiayuan Yang, Zhiyuan Zhao, Da Zhang 等CVPR 2026 · 被引用 3 次
它引用的顶会 Paper34
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
相关 Paper
- OmniSVG: A Unified Scalable Vector Graphics Generation ModelYiying Yang, Wei Cheng, Sijin Chen, Xianfang Zeng 等NeurIPS 2025 · 被引用 90 次
- SVGThinker: Instruction-Aligned and Reasoning-Driven Text-to-SVG GenerationHanqi Chen, Zhongyin Zhao, Ye Chen, Zhujin Liang 等ACM MM 2025 · 被引用 3 次
- Vector Prism: Animating Vector Graphics by Stratifying Semantic StructureJooyeol Yun, Jaegul ChooCVPR 2026 · 被引用 2 次
- Seeing is Improving: Visual Feedback for Iterative Text Layout RefinementJunrong Guo, Shancheng Fang, Yadong Qu, Hongtao XieCVPR 2026 · 被引用 2 次
- SVGen: Interpretable Vector Graphics Generation with Large Language ModelsFeiyu Wang, Zhiyuan Zhao, Yuandong Liu, Da Zhang 等ACM MM 2025 · 被引用 7 次
