RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills
Chunru Lin, Haotian Yuan, Yian Wang, Xiaowen Qiu, Tsun-Hsuan Johnson Wang, Minghao Guo, Bohan Wang, Yashraj Narang, Dieter Fox, Chuang Gan
摘要
Endowing robots with tool design abilities is critical for enabling them to solve complex manipulation tasks that would otherwise be intractable. While recent generative frameworks can automatically synthesize task settings, such as 3D scenes and reward functions, they have not yet addressed the challenge of tooluse scenarios. Simply retrieving human-designed tools might not be ideal since many tools (e.g., a rolling pin) are difficult for robotic manipulators to handle. Furthermore, existing tool design approaches either rely on predefined templates with limited parameter tuning or apply generic 3D generation methods that are not optimized for tool creation. To address these limitations, we propose RobotSmith, an automated pipeline that leverages the implicit physical knowledge embedded in vision-language models (VLMs) alongside the more accurate physics provided by physics simulations to design and use tools for robotic manipulation. Our system (1) iteratively proposes tool designs using collaborative VLM agents, (2) generates low-level robot trajectories for tool use, and (3) jointly optimizes tool geometry and usage for task performance. We evaluate our approach across a wide range of manipulation tasks involving rigid, deformable, and fluid objects. Experiments show that our method consistently outperforms strong baselines in terms of both task success rate and overall performance. Notably, our approach achieves a 50.0% average success rate, significantly surpassing other baselines such as 3D generation (21.4%) and tool retrieval (11.1%). Finally, we deploy our system in real-world settings, demonstrating that the generated tools and their usage plans transfer effectively to physical execution, validating the practicality and generalization capabilities of our approach. 2 Recently, a number of generative frameworks for robotic manipulation have emerged [54,51,9,53,12,20,23,32], demonstrating strong potential to scale up data collection and task generalization. However, these frameworks have overlooked tool-use scenarios, as they typically rely on retrieving * denotes equal contribution 2 Project page: https://umass-embodied-agi.github.io/RobotSmith/ 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- SpatialRGPT: Grounded Spatial Reasoning in Vision-Language ModelsAn-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo 等NeurIPS 2024 · 被引用 412 次
- RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative SimulationYufei Wang, Zhou Xian, Feng Chen, Tsun-Hsuan Wang 等ICML 2024 · 被引用 227 次
- LIV: Language-Image Representations and Rewards for Robotic ControlYecheng Jason Ma, Vikash Kumar, Amy Zhang, Osbert Bastani 等ICML 2023 · 被引用 212 次
- RoboCLIP: One Demonstration is Enough to Learn Robot PoliciesSumedh Sontakke, Jesse Zhang, Sébastien M. R. Arnold, Karl Pertsch 等NeurIPS 2023 · 被引用 182 次
- Vision-Language Models are Zero-Shot Reward Models for Reinforcement LearningJuan Rocamonde, Victoriano Montesinos, Elvis Nava, Ethan Perez 等ICLR 2024 · 被引用 154 次
相关 Paper
- VLMgineer: Vision-Language Models as Robotic ToolsmithsGeorge Jiayuan Gao, Tianyu Li, Junyao Shi, Yihan Li 等ICLR 2026 · 被引用 14 次
- SceneSmith: Agentic Generation of Simulation-Ready Indoor ScenesNicholas Pfaff, Thomas Cohn, Sergey Zakharov, Rick Cory 等ICML 2026 · 被引用 21 次
- SIMPACT: Simulation-Enabled Action Planning using Vision-Language ModelsHaowen Liu, Shaoxiong Yao, Haonan Chen, Jiawei Gao 等CVPR 2026 · 被引用 8 次
- ManipEvalAgent: Promptable and Efficient Evaluation Framework for Robotic Manipulation PoliciesYiteng Chen, Huiping Zhuang, Wenbo Li, Shiyi Wang 等ICLR 2026
- PAT3D: Physics-Augmented Text-to-3D Scene GenerationGuying Lin, Kemeng Huang, Michael Liu, Ruihan Gao 等ICLR 2026 · 被引用 14 次
