CMT: A Cascade MAR with Topology Predictor for Multimodal Conditional CAD Generation
Jianyu Wu, Yizhou Wang, Xiangyu Yue, Xinzhu Ma, Jinyang Guo, Dongzhan Zhou, Wanli Ouyang, Shixiang Tang
摘要
While accurate and user-friendly Computer-Aided Design (CAD) is crucial for industrial design and manufacturing, existing methods still struggle to achieve this due to their over-simplified representations or architectures incapable of supporting multimodal design requirements. In this paper, we attempt to tackle this problem from both methods and datasets aspects. First, we propose a cascade MAR with topology predictor (CMT), the first multimodal framework for CAD generation based on Boundary Representation (B-Rep). Specifically, the cascade MAR can effectively capture the ``edge-counters-surface'' priors that are essential in B-Reps, while the topology predictor directly estimates topology in B-Reps from the compact tokens in MAR. Second, to facilitate large-scale training, we develop a large-scale multimodal CAD dataset, mmABC, which includes over 1.3 million B-Rep models with multimodal annotations, including point clouds, text descriptions, and multi-view images. Extensive experiments show the superior of CMT in both conditional and unconditional CAD generation tasks. For example, we improve Coverage and Valid ratio by +10.68% and +10.3%, respectively, compared to state-of-the-art methods on ABC in unconditional generation. CMT also improves +4.01 Chamfer on image conditioned CAD generation on mmABC.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- MeshCoder: LLM-Powered Structured Mesh Code Generation from Point CloudsBingquan Dai, Li Ray Luo, Qihong Tang, Jie Wang 等NeurIPS 2025 · 被引用 19 次
- Pointer-CAD: Unifying B-Rep and Command Sequences via Pointer-based Edges & Faces SelectionDacheng Qi, Chenyu Wang, Jingwei Xu, Tianzhe Chu 等CVPR 2026 · 被引用 10 次
- CADFS: A Big CAD Program Dataset and Framework for Computer-Aided Design with Large Language ModelsVladislav Pyatov, Gleb Bobrovskikh, Saveliy Galochkin, Nikita Boldyrev 等CVPR 2026 · 被引用 4 次
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
相关 Paper
- Flatten the Complex: Joint B-Rep Generation via Compositional k-Cell ParticlesJunran Lu, Yuanqi Li, Hengji Li, Jie Guo 等SIGGRAPH 2026
- B-repLer: Language-guided Editing of CAD ModelsYilin Liu, Niladri Shekhar Dutt, Changjian Li, Niloy J. MitraSIGGRAPH 2026 · 被引用 1 次
- CAD Translator: An Effective Drive for Text to 3D Parametric Computer-Aided Design Generative ModelingXueyang Li, Yu Song, Yunzhong Lou, Xiangdong ZhouACM MM 2024 · 被引用 16 次
- FreeCAD: A Multimodal Framework for 3D CAD Model Generation from Free-Form PromptsDawei Lin, Meng Yuan, Ziming Wang, Tieru Wu 等ACM MM 2025 · 被引用 4 次
- ComplexGen: CAD reconstruction by B-rep chain complex generationHaoxiang Guo, Shilin Liu, Hao Pan, Yang Liu 等SIGGRAPH 2022 · 被引用 106 次
