ContextSeg: Sketch Semantic Segmentation by Querying the Context with Attention
Jiawei Wang, Changjian Li
Abstract
Sketch semantic segmentation is a well-explored and pivotal problem in computer vision involving the assignment of pre-defined part labels to individual strokes. This paper presents ContextSeg-a simple yet highly effective approach to tackling this problem with two stages. In the first stage, to better encode the shape and positional information of strokes, we propose to predict an extra dense distance field in an autoencoder network to reinforce structural information learning. In the second stage, we treat an entire stroke as a single entity and label a group of strokes within the same semantic part using an auto-regressive Transformer with the default attention mechanism. By group-based labeling, our method can fully leverage the context information when making decisions for the remaining groups of strokes. Our method achieves the best segmentation accuracy compared with state-of-the-art approaches on two representative datasets and has been extensively evaluated demonstrating its superior performance. Additionally, we offer insights into solving part imbalance in training data and the preliminary experiment on cross-category training, which can inspire future research in this field.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 57ef04a2-e287-45d9-a79d-5955a803af79Cited by top-tier papers3
- Instance Segmentation of Scene Sketches Using Natural Image PriorsMia Tang, Yael Vinker, Chuan Yan, Lvmin Zhang et al.SIGGRAPH 2025 · 3 citations
- VQ-SGen: A Vector Quantized Stroke Representation for Creative Sketch GenerationJiawei Wang, Zhiming Cui, Changjian LiICCV 2025 · 3 citations
- SketchAgent: Language-Driven Sequential Sketch GenerationYael Vinker, Tamar Rott Shaham, Kristine Zheng, Alex Zhao et al.CVPR 2025
Builds on6
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li et al.ICCV 2021 · 1,611 citations
- Free2CAD: parsing freehand drawings into CAD commandsChangjian Li, Hao Pan, Adrien Bousseau, Niloy J. MitraSIGGRAPH 2022 · 100 citations
- Creative Sketch GenerationSongwei Ge, Vedanuj Goswami, Larry Zitnick, Devi ParikhICLR 2021
Related papers
- Open Vocabulary Semantic Scene Sketch UnderstandingAhmed Bourouis, Judith Ellen Fan, Yulia GryaditskayaCVPR 2024
- Stroke2Sketch: Harnessing Stroke Attributes for Training-Free Sketch GenerationRui Yang, Huining Li, Yiyi Long, Xiaojun Wu et al.ICCV 2025 · 2 citations
- CoSE: Compositional Stroke EmbeddingsEmre Aksan, Thomas Deselaers, Andrea Tagliasacchi, Otmar HilligesNeurIPS 2020 · 37 citations
- Stroke Extraction of Chinese Character Based on Deep Structure Deformable Image RegistrationMeng Li, Yahan Yu, Yi Yang, Guanghao Ren et al.AAAI 2023 · 7 citations
- Generating Sketches in a Hierarchical Auto-Regressive Process for Flexible Sketch Drawing Manipulation at Stroke-LevelSicong Zang, Shuhui Gao, Zhijun FangAAAI 2026 · 2 citations
