ShapeAR: Generating Editable Shape Layers via Autoregressive Diffusion
Souymodip Chakraborty, Ankur Singh, Amit Vikram Singh, Vineet Batra, Ankit Phogat
Abstract
We present ShapeAR, a novel autoregressive latent diffusion framework that decomposes raster images into editable, artist-like vector shape layers. Unlike conventional raster-to-SVG methods that rely on boundary tracing or joint path optimization, ShapeAR generates non-overlapping RGBA shape layers directly in latent space via flow-matching diffusion. To scale generation to complex scenes with many shapes, we formulate the process autoregressively, conditioning each step on both the input image (global context) and the partial composition of previously generated layers (local context). In addition, we propose geometry-aware evaluation metrics that quantify the aesthetic and structural quality of the generated shapes, enabling more rigorous assessment beyond pixel-level reconstruction. ShapeAR achieves cleaner decompositions and more coherent vector layers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0b609201-b9d4-4298-a88a-b2ca35eb933dBuilds on19
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- CLIPDraw: Exploring Text-to-Drawing Synthesis through Language-Image EncodersKevin Frans, Lisa B. Soros, Olaf WitkowskiNeurIPS 2022 · 311 citations
- DeepSVG: A Hierarchical Generative Network for Vector Graphics AnimationAlexandre Carlier, Martin Danelljan, Alexandre Alahi, Radu TimofteNeurIPS 2020 · 247 citations
- A Learned Representation for Scalable Vector GraphicsRaphael Gontijo Lopes, David Ha, Douglas Eck, Jonathon ShlensICCV 2019 · 153 citations
Related papers
- LayerTracer: Cognitive-Aligned Layered SVG Synthesis via Diffusion TransformerYiren Song, Danze Chen, Mike Zheng ShouICCV 2025 · 5 citations
- Text-to-Vector Generation with Neural Path RepresentationPeiying Zhang, Nanxuan Zhao, Jing LiaoSIGGRAPH 2024 · 15 citations
- DesignEdit: Unify Spatial-Aware Image Editing via Training-free Inpainting with a Multi-Layered Latent Diffusion FrameworkYueru Jia, Aosong Cheng, Yuhui Yuan, Chuke Wang et al.AAAI 2025 · 5 citations
- SwiftSketch: A Diffusion Model for Image-to-Vector Sketch GenerationEllie Arar, Yarden Frenkel, Daniel Cohen-Or, Ariel Shamir et al.SIGGRAPH 2025 · 12 citations
- ControlAR: Controllable Image Generation with Autoregressive ModelsZongming Li, Tianheng Cheng, Shoufa Chen, Peize Sun et al.ICLR 2025
