VLMgineer: Vision-Language Models as Robotic Toolsmiths
George Jiayuan Gao, Tianyu Li, Junyao Shi, Yihan Li, Zizhe Zhang, Nadia Figueroa, Dinesh Jayaraman
Abstract
Tool design and use reflect the ability to understand and manipulate the physical world through creativity, planning, and foresight. As such, it is often regarded as a measurable indicator of cognitive intelligence across biological species. While much of today’s research on robotics intelligence focuses on generating better control strategies, inventing smarter tools offers a complementary form of physical intelligence: moving the problem-solving onus into the tool’s geometry so that control becomes simpler. This motivates us to ask: can today’s foundation models offer useful priors to automatically invent—and effectively wield—such tools? We present VLMgineer, the first fully automatic framework designs tools and actions from scratch by harnessing the creativity of Vision–Language Models (VLMs) together with evolutionary search. We evaluate VLMgineer on a diverse benchmark of everyday manipulation scenarios that demand creative tool design and use. Across this suite, VLMgineer consistently discovers tools and policies that solve tasks more effectively and innovatively, transforming challenging robotics problems into straightforward executions. It also consistently outperforms VLM-generated designs from human specifications and existing human-crafted tools for everyday tasks. We further demonstrate that VLMgineer’s automatically designed tools and action policies transfer seamlessly to real-world task execution on a physical robot. To facilitate future research on automated tool invention, we will release our benchmark and code. Project Website: https://vlmgineer.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c3bce744-bef4-4d4e-98e3-4c460ad6a9f7Cited by top-tier papers1
Ask how each one uses itBuilds on12
- Eureka: Human-Level Reward Design via Coding Large Language ModelsYecheng Jason Ma, William Liang, Guanzhi Wang, De-An Huang et al.ICLR 2024 · 582 citations
- Evolution Gym: A Large-Scale Benchmark for Evolving Soft RobotsJagdeep Singh Bhatia, Holly Jackson, Yunsheng Tian, Jie Xu et al.NeurIPS 2021 · 141 citations
- DiffSkill: Skill Abstraction from Differentiable Physics for Deformable Object Manipulations with ToolsXingyu Lin, Zhiao Huang, Yunzhu Li, Joshua B. Tenenbaum et al.ICLR 2022 · 85 citations
- Transform2Act: Learning a Transform-and-Control Policy for Efficient Agent DesignYe Yuan, Yuda Song, Zhengyi Luo, Wen Sun et al.ICLR 2022 · 51 citations
- Task-Agnostic Morphology EvolutionDonald Joseph Hejna III, Pieter Abbeel, Lerrel PintoICLR 2021 · 32 citations
Related papers
- RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation SkillsChunru Lin, Haotian Yuan, Yian Wang, Xiaowen Qiu et al.NeurIPS 2025 · 10 citations
- Transductive Visual Programming: Evolving Tool Libraries from Experience for Spatial ReasoningShengguang Wu, Xiaohan Wang, Yuhui Zhang, Hao Zhu et al.ICLR 2026 · 2 citations
- Vision-Language Foundation Models as Effective Robot ImitatorsXinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu et al.ICLR 2024 · 375 citations
- LLaRA: Supercharging Robot Learning Data for Vision-Language PolicyXiang Li, Cristina Mata, Jongwoo Park, Kumara Kahatapitiya et al.ICLR 2025 · 2 citations
- CAD-Assistant: Tool-Augmented VLLMs as Generic CAD Task SolversDimitrios Mallis, Ahmet Serdar Karadeniz, Sebastian Cavada, Danila Rukhovich et al.ICCV 2025 · 9 citations
