Rethinking Human Intent-to-CAD: Parametric CAD Model Generation via Cooperative Multi-Task Alignment and Spatial-Aware Reinforcement Learning
Qingwang Zhang, Jiahao Li, Xiangdong Zhou
摘要
Parametric Computer-Aided-Design (CAD) modeling from human intent remains challenging, particularly during the conceptual design stage, where design goals are expressed through incomplete and unstructured modalities (e.g., handdrawn sketches and textual descriptions). In this work, we rethink the human intent-to-CAD pipeline and propose a unified method that directly maps multi-level human intents to executable codes, without assuming the prior existence of target CAD models. To support our study, we construct HiCAD, the first large-scale dataset aligning hand-drawn sketches, textual descriptions, and parametric CAD codes. Based on this, we introduce HiCAD, a two-stage framework comprising Cooperative Multi-Task Alignment to bridge the representational gap between heterogeneous inputs, and Spatial-Aware Reinforcement Learning to enforce geometric and topological consistency. Extensive experiments demonstrate that our method significantly outperforms existing baselines across multiple tasks, validating its effectiveness and robustness in transforming heterogeneous human intents into high-fidelity parametric CAD models. Our project page: https: //zqwlearning.github.io/HiCAD.
You will be given a CADQuery code snippet <cadquery> and a rendered image <image>. The CAD model has a length (Xaxis) of <length> units, width (Y-axis) of <width> units, height (Z-axis) of <height> units, and includes <through-holes> through holes.
You are a senior CAD engineer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
相关 Paper
- From Intent to Execution: Multimodal Chain-of-Thought Reinforcement Learning for Precise CAD Code GenerationKe Niu, Haiyang Yu, Zhuofan Chen, Mengyang Zhao 等AAAI 2026 · 被引用 5 次
- ReCAD: Reinforcement Learning Enhanced Parametric CAD Model Generation with Vision-Language ModelsJiahao Li, Yusheng Luo, Yunzhong Lou, Xiangdong ZhouAAAI 2026 · 被引用 4 次
- CAD-Coder: Text-to-CAD Generation with Chain-of-Thought and Geometric RewardYandong Guan, Xilin Wang, Ximing Xing, Jing Zhang 等NeurIPS 2025 · 被引用 64 次
- CME-CAD: Heterogeneous Collaborative Multi-Expert Reinforcement Learning for CAD Code GenerationKe Niu, Haiyang Yu, Zhuofan Chen, Zhengtao Yao 等CVPR 2026 · 被引用 18 次
- CAD-GPT: Synthesising CAD Construction Sequence with Spatial Reasoning-Enhanced Multimodal LLMsSiyu Wang, Cailian Chen, Xinyi Le, Qimin Xu 等AAAI 2025 · 被引用 49 次
