Text-to-CAD Generation Through Infusing Visual Feedback in Large Language Models
Ruiyu Wang, Yu Yuan, Shizhao Sun, Jiang Bian
摘要
Creating Computer-Aided Design (CAD) models requires considerable expertise and effort. Textto-CAD, which converts textual descriptions into CAD parametric sequences, is crucial in streamlining this process. Recent studies have utilized ground-truth parametric sequences, known as sequential signals, as supervision to achieve this goal. However, CAD models are inherently multimodal, comprising parametric sequences and corresponding rendered visual objects. Besides, the rendering process from parametric sequences to visual objects is many-to-one. Therefore, both sequential and visual signals are critical for effective training. In this work, we introduce CADFusion, a framework that uses Large Language Models (LLMs) as the backbone and alternates between two training stages: the sequential learning (SL) stage and the visual feedback (VF) stage. In the SL stage, we train LLMs using ground-truth parametric sequences, enabling the generation of logically coherent parametric sequences. In the VF stage, we reward parametric sequences that render into visually preferred objects and penalize those that do not, allowing LLMs to learn how rendered visual objects are perceived and evaluated. These two stages alternate throughout the training, ensuring balanced learning and preserving benefits of both signals. Experiments demonstrate that CAD-Fusion improves performance, both qualitatively and quantitatively. Code is available at https: //github.com/microsoft/CADFusion .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- CAD-Coder: Text-to-CAD Generation with Chain-of-Thought and Geometric RewardYandong Guan, Xilin Wang, Ximing Xing, Jing Zhang 等NeurIPS 2025 · 被引用 64 次
- Rendering-Aware Reinforcement Learning for Vector Graphics GenerationJuan A. Rodríguez, Haotian Zhang, Abhay Puri, Rishav Pramanik 等NeurIPS 2025 · 被引用 42 次
- UORA: Uniform Orthogonal Reinitialization Adaptation in Parameter Efficient Fine-Tuning of Large ModelsXueyan Zhang, Jinman Zhao, Zhifei Yang, Yibo Zhong 等ACL 2025 · 被引用 15 次
- Pointer-CAD: Unifying B-Rep and Command Sequences via Pointer-based Edges & Faces SelectionDacheng Qi, Chenyu Wang, Jingwei Xu, Tianzhe Chu 等CVPR 2026 · 被引用 10 次
- Automated CAD Modeling Sequence Generation from Text Descriptions via Transformer-Based Large Language ModelsJianxing Liao, Junyan Xu, Yatao Sun, Maowen Tang 等ACL 2025 · 被引用 8 次
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
- DeepCAD: A Deep Generative Network for Computer-Aided Design ModelsRundi Wu, Chang Xiao, Changxi ZhengICCV 2021 · 被引用 290 次
相关 Paper
- FreeCAD: A Multimodal Framework for 3D CAD Model Generation from Free-Form PromptsDawei Lin, Meng Yuan, Ziming Wang, Tieru Wu 等ACM MM 2025 · 被引用 4 次
- Multi-Agent CAD Code GenerationYang Liu, Daxuan Ren, Yijie Ding, Jianmin Zheng 等SIGGRAPH 2026
- CAD Translator: An Effective Drive for Text to 3D Parametric Computer-Aided Design Generative ModelingXueyang Li, Yu Song, Yunzhong Lou, Xiangdong ZhouACM MM 2024 · 被引用 16 次
- ReCAD: Reinforcement Learning Enhanced Parametric CAD Model Generation with Vision-Language ModelsJiahao Li, Yusheng Luo, Yunzhong Lou, Xiangdong ZhouAAAI 2026 · 被引用 4 次
- Towards High-Fidelity CAD Generation via LLM-Driven Program Generation and Text-Based B-Rep Primitive GroundingJiahao Li, Qingwang Zhang, Qiuyu Chen, Guozhan Qiu 等ICML 2026 · 被引用 7 次
