CADTransformer: Panoptic Symbol Spotting Transformer for CAD Drawings
Zhiwen Fan, Tianlong Chen, Peihao Wang, Zhangyang Wang
摘要
Understanding 2D computer-aided design (CAD) drawings plays a crucial role for creating 3D prototypes in architecture, engineering and construction (AEC) industries. The task of automated panoptic symbol spotting, i.e., to spot and parse both countable object instances (windows, doors, tables, etc.) and uncountable stuff (wall, railing, etc.) from CAD drawings, has recently drawn interests from the computer vision community. Unfortunately, the highly irregular ordering and orientations set major roadblocks for this task. Existing methods, based on convolutional neural networks (CNNs) and/or graph neural networks (GNNs), regress instance bounding boxes in the pixel domain and then convert the predictions into symbols. In this paper, we present a novel framework named CADTransformer, that can painlessly modify existing vision transformer (ViT) backbones to tackle the above limitations for the panoptic symbol spotting task. CADTransformer tokenizes directly from the set of graphical primitives in CAD drawings, and correspondingly optimizes line-grained semantic and instance symbol spotting altogether by a pair of prediction heads. The backbone is further enhanced with a few plug-and-play modifications, including a neighborhood aware self-attention, hierarchical feature aggregation, and graphic entity position encoding, to bake in the structure prior while optimizing the efficiency. Besides, a new data augmentation method, termed Random Layer, is proposed by the layer-wise separation and recombination of a CAD drawing. Overall, CADTransformer significantly boosts the previous state-of-the-art from 0.595 to 0.685 in the panoptic quality (PQ) metric, on the recently released FloorPlanCAD dataset. We further demonstrate that our model can spot symbols with irregular shapes and arbitrary orientations. Our codes are available in https: //github.com/VITA-Group/CADTransformer .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Simplifying and Empowering Transformers for Large-Graph RepresentationsQitian Wu, Wentao Zhao, Chenxiao Yang, Hengrui Zhang 等NeurIPS 2023 · 被引用 318 次
- PlankAssembly: Robust 3D Reconstruction from Three Orthographic Views with Learnt Shape ProgramsWentao Hu, Jia Zheng, Zixin Zhang, Xiaojun Yuan 等ICCV 2023 · 被引用 12 次
- Symbol as Points: Panoptic Symbol Spotting via Point-based RepresentationWenlong Liu, Tianyu Yang, Yuhan Wang, Qizhi Yu 等ICLR 2024 · 被引用 10 次
- Symbolic Distillation for Learned TCP Congestion ControlS. P. Sharan, Wenqing Zheng, Kuo-Feng Hsu, Jiarong Xing 等NeurIPS 2022 · 被引用 9 次
- ArchCAD-400K: A Large-Scale CAD drawings Dataset and New Baseline for Panoptic Symbol SpottingRuifeng Luo, Zhengjie Liu, Tianxiao Cheng, Jie Wang 等NeurIPS 2025 · 被引用 8 次
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
相关 Paper
- FloorPlanCAD: A Large-Scale CAD Drawing Dataset for Panoptic Symbol SpottingZhiwen Fan, Lingjie Zhu, Honghua Li, Xiaohao Chen 等ICCV 2021 · 被引用 53 次
- GAT-CADNet: Graph Attention Network for Panoptic Symbol Spotting in CAD DrawingsZhaohua Zheng, Jianfang Li, Lingjie Zhu, Honghua Li 等CVPR 2022 · 被引用 19 次
- Point or Line? Using Line-based Representation for Panoptic Symbol Spotting in CAD DrawingsXingguang Wei, Haomin Wang, Shenglong Ye, Ruifeng Luo 等NeurIPS 2025 · 被引用 5 次
- Free2CAD: parsing freehand drawings into CAD commandsChangjian Li, Hao Pan, Adrien Bousseau, Niloy J. MitraSIGGRAPH 2022 · 被引用 100 次
- DeepCAD: A Deep Generative Network for Computer-Aided Design ModelsRundi Wu, Chang Xiao, Changxi ZhengICCV 2021 · 被引用 290 次
