CADTransformer: Panoptic Symbol Spotting Transformer for CAD Drawings
Zhiwen Fan, Tianlong Chen, Peihao Wang, Zhangyang Wang
Abstract
Understanding 2D computer-aided design (CAD) drawings plays a crucial role for creating 3D prototypes in architecture, engineering and construction (AEC) industries. The task of automated panoptic symbol spotting, i.e., to spot and parse both countable object instances (windows, doors, tables, etc.) and uncountable stuff (wall, railing, etc.) from CAD drawings, has recently drawn interests from the computer vision community. Unfortunately, the highly irregular ordering and orientations set major roadblocks for this task. Existing methods, based on convolutional neural networks (CNNs) and/or graph neural networks (GNNs), regress instance bounding boxes in the pixel domain and then convert the predictions into symbols. In this paper, we present a novel framework named CADTransformer, that can painlessly modify existing vision transformer (ViT) backbones to tackle the above limitations for the panoptic symbol spotting task. CADTransformer tokenizes directly from the set of graphical primitives in CAD drawings, and correspondingly optimizes line-grained semantic and instance symbol spotting altogether by a pair of prediction heads. The backbone is further enhanced with a few plug-and-play modifications, including a neighborhood aware self-attention, hierarchical feature aggregation, and graphic entity position encoding, to bake in the structure prior while optimizing the efficiency. Besides, a new data augmentation method, termed Random Layer, is proposed by the layer-wise separation and recombination of a CAD drawing. Overall, CADTransformer significantly boosts the previous state-of-the-art from 0.595 to 0.685 in the panoptic quality (PQ) metric, on the recently released FloorPlanCAD dataset. We further demonstrate that our model can spot symbols with irregular shapes and arbitrary orientations. Our codes are available in https: //github.com/VITA-Group/CADTransformer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d19872c9-5e4a-4762-a6a9-a0cc253b1615Cited by top-tier papers13
- Simplifying and Empowering Transformers for Large-Graph RepresentationsQitian Wu, Wentao Zhao, Chenxiao Yang, Hengrui Zhang et al.NeurIPS 2023 · 318 citations
- PlankAssembly: Robust 3D Reconstruction from Three Orthographic Views with Learnt Shape ProgramsWentao Hu, Jia Zheng, Zixin Zhang, Xiaojun Yuan et al.ICCV 2023 · 12 citations
- Symbol as Points: Panoptic Symbol Spotting via Point-based RepresentationWenlong Liu, Tianyu Yang, Yuhan Wang, Qizhi Yu et al.ICLR 2024 · 10 citations
- Symbolic Distillation for Learned TCP Congestion ControlS. P. Sharan, Wenqing Zheng, Kuo-Feng Hsu, Jiarong Xing et al.NeurIPS 2022 · 9 citations
- ArchCAD-400K: A Large-Scale CAD drawings Dataset and New Baseline for Panoptic Symbol SpottingRuifeng Luo, Zhengjie Liu, Tianxiao Cheng, Jie Wang et al.NeurIPS 2025 · 8 citations
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
Related papers
- FloorPlanCAD: A Large-Scale CAD Drawing Dataset for Panoptic Symbol SpottingZhiwen Fan, Lingjie Zhu, Honghua Li, Xiaohao Chen et al.ICCV 2021 · 53 citations
- GAT-CADNet: Graph Attention Network for Panoptic Symbol Spotting in CAD DrawingsZhaohua Zheng, Jianfang Li, Lingjie Zhu, Honghua Li et al.CVPR 2022 · 19 citations
- Point or Line? Using Line-based Representation for Panoptic Symbol Spotting in CAD DrawingsXingguang Wei, Haomin Wang, Shenglong Ye, Ruifeng Luo et al.NeurIPS 2025 · 5 citations
- Free2CAD: parsing freehand drawings into CAD commandsChangjian Li, Hao Pan, Adrien Bousseau, Niloy J. MitraSIGGRAPH 2022 · 100 citations
- DeepCAD: A Deep Generative Network for Computer-Aided Design ModelsRundi Wu, Chang Xiao, Changxi ZhengICCV 2021 · 290 citations
