ShapeAR: Generating Editable Shape Layers via Autoregressive Diffusion
Souymodip Chakraborty, Ankur Singh, Amit Vikram Singh, Vineet Batra, Ankit Phogat
摘要
We present ShapeAR, a novel autoregressive latent diffusion framework that decomposes raster images into editable, artist-like vector shape layers. Unlike conventional raster-to-SVG methods that rely on boundary tracing or joint path optimization, ShapeAR generates non-overlapping RGBA shape layers directly in latent space via flow-matching diffusion. To scale generation to complex scenes with many shapes, we formulate the process autoregressively, conditioning each step on both the input image (global context) and the partial composition of previously generated layers (local context). In addition, we propose geometry-aware evaluation metrics that quantify the aesthetic and structural quality of the generated shapes, enabling more rigorous assessment beyond pixel-level reconstruction. ShapeAR achieves cleaner decompositions and more coherent vector layers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- CLIPDraw: Exploring Text-to-Drawing Synthesis through Language-Image EncodersKevin Frans, Lisa B. Soros, Olaf WitkowskiNeurIPS 2022 · 被引用 311 次
- DeepSVG: A Hierarchical Generative Network for Vector Graphics AnimationAlexandre Carlier, Martin Danelljan, Alexandre Alahi, Radu TimofteNeurIPS 2020 · 被引用 247 次
- A Learned Representation for Scalable Vector GraphicsRaphael Gontijo Lopes, David Ha, Douglas Eck, Jonathon ShlensICCV 2019 · 被引用 153 次
相关 Paper
- LayerTracer: Cognitive-Aligned Layered SVG Synthesis via Diffusion TransformerYiren Song, Danze Chen, Mike Zheng ShouICCV 2025 · 被引用 5 次
- Text-to-Vector Generation with Neural Path RepresentationPeiying Zhang, Nanxuan Zhao, Jing LiaoSIGGRAPH 2024 · 被引用 15 次
- DesignEdit: Unify Spatial-Aware Image Editing via Training-free Inpainting with a Multi-Layered Latent Diffusion FrameworkYueru Jia, Aosong Cheng, Yuhui Yuan, Chuke Wang 等AAAI 2025 · 被引用 5 次
- SwiftSketch: A Diffusion Model for Image-to-Vector Sketch GenerationEllie Arar, Yarden Frenkel, Daniel Cohen-Or, Ariel Shamir 等SIGGRAPH 2025 · 被引用 12 次
- ControlAR: Controllable Image Generation with Autoregressive ModelsZongming Li, Tianheng Cheng, Shoufa Chen, Peize Sun 等ICLR 2025
