VLMgineer: Vision-Language Models as Robotic Toolsmiths
George Jiayuan Gao, Tianyu Li, Junyao Shi, Yihan Li, Zizhe Zhang, Nadia Figueroa, Dinesh Jayaraman
摘要
Tool design and use reflect the ability to understand and manipulate the physical world through creativity, planning, and foresight. As such, it is often regarded as a measurable indicator of cognitive intelligence across biological species. While much of today’s research on robotics intelligence focuses on generating better control strategies, inventing smarter tools offers a complementary form of physical intelligence: moving the problem-solving onus into the tool’s geometry so that control becomes simpler. This motivates us to ask: can today’s foundation models offer useful priors to automatically invent—and effectively wield—such tools? We present VLMgineer, the first fully automatic framework designs tools and actions from scratch by harnessing the creativity of Vision–Language Models (VLMs) together with evolutionary search. We evaluate VLMgineer on a diverse benchmark of everyday manipulation scenarios that demand creative tool design and use. Across this suite, VLMgineer consistently discovers tools and policies that solve tasks more effectively and innovatively, transforming challenging robotics problems into straightforward executions. It also consistently outperforms VLM-generated designs from human specifications and existing human-crafted tools for everyday tasks. We further demonstrate that VLMgineer’s automatically designed tools and action policies transfer seamlessly to real-world task execution on a physical robot. To facilitate future research on automated tool invention, we will release our benchmark and code. Project Website: https://vlmgineer.github.io/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper12
- Eureka: Human-Level Reward Design via Coding Large Language ModelsYecheng Jason Ma, William Liang, Guanzhi Wang, De-An Huang 等ICLR 2024 · 被引用 582 次
- Evolution Gym: A Large-Scale Benchmark for Evolving Soft RobotsJagdeep Singh Bhatia, Holly Jackson, Yunsheng Tian, Jie Xu 等NeurIPS 2021 · 被引用 141 次
- DiffSkill: Skill Abstraction from Differentiable Physics for Deformable Object Manipulations with ToolsXingyu Lin, Zhiao Huang, Yunzhu Li, Joshua B. Tenenbaum 等ICLR 2022 · 被引用 85 次
- Transform2Act: Learning a Transform-and-Control Policy for Efficient Agent DesignYe Yuan, Yuda Song, Zhengyi Luo, Wen Sun 等ICLR 2022 · 被引用 51 次
- Task-Agnostic Morphology EvolutionDonald Joseph Hejna III, Pieter Abbeel, Lerrel PintoICLR 2021 · 被引用 32 次
相关 Paper
- RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation SkillsChunru Lin, Haotian Yuan, Yian Wang, Xiaowen Qiu 等NeurIPS 2025 · 被引用 10 次
- Transductive Visual Programming: Evolving Tool Libraries from Experience for Spatial ReasoningShengguang Wu, Xiaohan Wang, Yuhui Zhang, Hao Zhu 等ICLR 2026 · 被引用 2 次
- Vision-Language Foundation Models as Effective Robot ImitatorsXinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu 等ICLR 2024 · 被引用 375 次
- LLaRA: Supercharging Robot Learning Data for Vision-Language PolicyXiang Li, Cristina Mata, Jongwoo Park, Kumara Kahatapitiya 等ICLR 2025 · 被引用 2 次
- CAD-Assistant: Tool-Augmented VLLMs as Generic CAD Task SolversDimitrios Mallis, Ahmet Serdar Karadeniz, Sebastian Cavada, Danila Rukhovich 等ICCV 2025 · 被引用 9 次
