CAD-Tokenizer: Towards Text-Based CAD Prototyping via Modality-Specific Tokenization
Ruiyu Wang, Shizhao Sun, Weijian Ma, Jiang Bian
摘要
Computer-Aided Design (CAD) is a foundational component of industrial prototyping. where models are defined not by raw coordinates but by construction sequences such as sketches and extrusions. This sequential structure enables both efficient prototype initialization and subsequent editing. Text-guided CAD prototyping, which unifies Text-to-CAD generation and CAD editing, has the potential to streamline the entire design pipeline. However, prior work has not explored this setting, largely because standard large language model (LLM) tokenizers decompose CAD sequences into natural-language word pieces, failing to capture primitive-level CAD semantics and hindering attention modules from modeling geometric structure. We conjecture that a multimodal tokenization strategy, aligned with CAD’s primitive and structural nature, can provide more effective representations. To this end, we propose CAD-Tokenizer, a framework that represents CAD data with modality-specific tokens using a sequence-based VQ-VAE with primitive-level pooling and constrained decoding. This design produces compact, primitive-aware representations that align with CAD’s structural nature. Applied to unified text-guided CAD prototyping, CAD-Tokenizer significantly improves instruction following and generation quality, achieving better quantitative and qualitative performance over both general-purpose LLMs and task-specific baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
相关 Paper
- Draw Step by Step: Reconstructing CAD Construction Sequences from Point Clouds via Multimodal DiffusionWeijian Ma, Shuaiqi Chen, Yunzhong Lou, Xueyang Li 等CVPR 2024 · 被引用 14 次
- Automated CAD Modeling Sequence Generation from Text Descriptions via Transformer-Based Large Language ModelsJianxing Liao, Junyan Xu, Yatao Sun, Maowen Tang 等ACL 2025 · 被引用 8 次
- FlexCAD: Unified and Versatile Controllable CAD Generation with Fine-tuned Large Language ModelsZhanwei Zhang, Shizhao Sun, Wenxiao Wang, Deng Cai 等ICLR 2025
- Pointer-CAD: Unifying B-Rep and Command Sequences via Pointer-based Edges & Faces SelectionDacheng Qi, Chenyu Wang, Jingwei Xu, Tianzhe Chu 等CVPR 2026 · 被引用 10 次
- CAD Translator: An Effective Drive for Text to 3D Parametric Computer-Aided Design Generative ModelingXueyang Li, Yu Song, Yunzhong Lou, Xiangdong ZhouACM MM 2024 · 被引用 16 次
