Text-to-CAD Generation Through Infusing Visual Feedback in Large Language Models
Ruiyu Wang, Yu Yuan, Shizhao Sun, Jiang Bian
Abstract
Creating Computer-Aided Design (CAD) models requires considerable expertise and effort. Textto-CAD, which converts textual descriptions into CAD parametric sequences, is crucial in streamlining this process. Recent studies have utilized ground-truth parametric sequences, known as sequential signals, as supervision to achieve this goal. However, CAD models are inherently multimodal, comprising parametric sequences and corresponding rendered visual objects. Besides, the rendering process from parametric sequences to visual objects is many-to-one. Therefore, both sequential and visual signals are critical for effective training. In this work, we introduce CADFusion, a framework that uses Large Language Models (LLMs) as the backbone and alternates between two training stages: the sequential learning (SL) stage and the visual feedback (VF) stage. In the SL stage, we train LLMs using ground-truth parametric sequences, enabling the generation of logically coherent parametric sequences. In the VF stage, we reward parametric sequences that render into visually preferred objects and penalize those that do not, allowing LLMs to learn how rendered visual objects are perceived and evaluated. These two stages alternate throughout the training, ensuring balanced learning and preserving benefits of both signals. Experiments demonstrate that CAD-Fusion improves performance, both qualitatively and quantitatively. Code is available at https: //github.com/microsoft/CADFusion .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7f60e539-343b-4c81-971d-405c95ac1602Cited by top-tier papers22
- CAD-Coder: Text-to-CAD Generation with Chain-of-Thought and Geometric RewardYandong Guan, Xilin Wang, Ximing Xing, Jing Zhang et al.NeurIPS 2025 · 64 citations
- Rendering-Aware Reinforcement Learning for Vector Graphics GenerationJuan A. Rodríguez, Haotian Zhang, Abhay Puri, Rishav Pramanik et al.NeurIPS 2025 · 42 citations
- UORA: Uniform Orthogonal Reinitialization Adaptation in Parameter Efficient Fine-Tuning of Large ModelsXueyan Zhang, Jinman Zhao, Zhifei Yang, Yibo Zhong et al.ACL 2025 · 15 citations
- Pointer-CAD: Unifying B-Rep and Command Sequences via Pointer-based Edges & Faces SelectionDacheng Qi, Chenyu Wang, Jingwei Xu, Tianzhe Chu et al.CVPR 2026 · 10 citations
- Automated CAD Modeling Sequence Generation from Text Descriptions via Transformer-Based Large Language ModelsJianxing Liao, Junyan Xu, Yatao Sun, Maowen Tang et al.ACL 2025 · 8 citations
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
- DeepCAD: A Deep Generative Network for Computer-Aided Design ModelsRundi Wu, Chang Xiao, Changxi ZhengICCV 2021 · 290 citations
Related papers
- FreeCAD: A Multimodal Framework for 3D CAD Model Generation from Free-Form PromptsDawei Lin, Meng Yuan, Ziming Wang, Tieru Wu et al.ACM MM 2025 · 4 citations
- Multi-Agent CAD Code GenerationYang Liu, Daxuan Ren, Yijie Ding, Jianmin Zheng et al.SIGGRAPH 2026
- CAD Translator: An Effective Drive for Text to 3D Parametric Computer-Aided Design Generative ModelingXueyang Li, Yu Song, Yunzhong Lou, Xiangdong ZhouACM MM 2024 · 16 citations
- ReCAD: Reinforcement Learning Enhanced Parametric CAD Model Generation with Vision-Language ModelsJiahao Li, Yusheng Luo, Yunzhong Lou, Xiangdong ZhouAAAI 2026 · 4 citations
- Towards High-Fidelity CAD Generation via LLM-Driven Program Generation and Text-Based B-Rep Primitive GroundingJiahao Li, Qingwang Zhang, Qiuyu Chen, Guozhan Qiu et al.ICML 2026 · 7 citations
