ReCAD: Reinforcement Learning Enhanced Parametric CAD Model Generation with Vision-Language Models
Jiahao Li, Yusheng Luo, Yunzhong Lou, Xiangdong Zhou
摘要
We present ReCAD, a reinforcement learning (RL) framework that bootstraps pretrained large models (PLMs) to generate precise parametric computer-aided design (CAD) models from multimodal inputs by leveraging their inherent generative capabilities. With just access to simple functional interfaces (e.g., point coordinates), our approach enables the emergence of complex CAD operations (e.g., pattern replication and mirror). This stands in contrast to previous methods, which typically rely on knowledge injected through supervised fine-tuning (SFT), offer limited support for editability, and fail to exploit the strong generative priors of PLMs. Specifically, the ReCAD framework begins by fine-tuning vision-language models (VLMs) to equip them with basic CAD model generation capabilities, where we rewrite CAD scripts into parameterized code that is leveraged to generate accurate textual descriptions for supervision. Then, we propose a novel RL strategy that incorporates parameterized code as guidance to enhance the model’s reasoning on challenging questions. Furthermore, we employ a hierarchical primitive learning process to progressively teach structured and compositional skills under a unified reward function that ensures both geometric accuracy and semantic fidelity. ReCAD sets a new state-of-the-art in both text-to-CAD and image-to-CAD tasks, significantly improving geometric accuracy across in-distribution and out-of-distribution settings. In the image-to-CAD task, for instance, it reduces the mean Chamfer Distance from 73.47 to 29.61 (in-distribution) and from 272.06 to 80.23 (out-of-distribution), outperforming existing baselines by a substantial margin.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Seek-CAD: A Self-refined Generative Modeling for 3D Parametric CAD Using Local Inference via DeepSeekXueyang Li, Jiahao Li, Yu Song, Yunzhong Lou 等ICLR 2026 · 被引用 30 次
- Towards High-Fidelity CAD Generation via LLM-Driven Program Generation and Text-Based B-Rep Primitive GroundingJiahao Li, Qingwang Zhang, Qiuyu Chen, Guozhan Qiu 等ICML 2026 · 被引用 7 次
- GIFT: Bootstrapping Image-to-CAD Program Synthesis via Geometric FeedbackGiorgio Giannone, Anna Doris, Amin Nobari, Kai Xu 等ICML 2026 · 被引用 3 次
- Rethinking Human Intent-to-CAD: Parametric CAD Model Generation via Cooperative Multi-Task Alignment and Spatial-Aware Reinforcement LearningQingwang Zhang, Jiahao Li, Xiangdong ZhouICML 2026
- Op-CAD: Benchmarking and Investigating Operation-oriented CAD GenerationYixue Bai, Yufei Gu, Zeke XieICML 2026
它引用的顶会 Paper23
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng 等EMNLP 2024 · 被引用 479 次
- Learning to Reason under Off-Policy GuidanceJianhao Yan, Yafu Li, Zican Hu, Zhi Wang 等NeurIPS 2025 · 被引用 310 次
- DeepCAD: A Deep Generative Network for Computer-Aided Design ModelsRundi Wu, Chang Xiao, Changxi ZhengICCV 2021 · 被引用 290 次
- Fusion 360 gallery: a dataset and environment for programmatic CAD construction from human design sequencesKarl D. D. Willis, Yewen Pu, Jieliang Luo, Hang Chu 等SIGGRAPH 2021 · 被引用 197 次
相关 Paper
- From Intent to Execution: Multimodal Chain-of-Thought Reinforcement Learning for Precise CAD Code GenerationKe Niu, Haiyang Yu, Zhuofan Chen, Mengyang Zhao 等AAAI 2026 · 被引用 5 次
- CADReview: Automatically Reviewing CAD Programs with Error Detection and CorrectionJiali Chen, Xusen Hei, Hongfei Liu, Yuancheng Wei 等ACL 2025
- CAD-Coder: Text-to-CAD Generation with Chain-of-Thought and Geometric RewardYandong Guan, Xilin Wang, Ximing Xing, Jing Zhang 等NeurIPS 2025 · 被引用 64 次
- CADFS: A Big CAD Program Dataset and Framework for Computer-Aided Design with Large Language ModelsVladislav Pyatov, Gleb Bobrovskikh, Saveliy Galochkin, Nikita Boldyrev 等CVPR 2026 · 被引用 4 次
- FreeCAD: A Multimodal Framework for 3D CAD Model Generation from Free-Form PromptsDawei Lin, Meng Yuan, Ziming Wang, Tieru Wu 等ACM MM 2025 · 被引用 4 次
