Scaffolding Dexterous Manipulation with Vision-Language Models
Vincent de Bakker, Joey Hejna, Tyler Ga Wei Lum, Onur Celik, Aleksandar Taranovic, Denis Blessing, Gerhard Neumann, Jeannette Bohg, Dorsa Sadigh
摘要
Dexterous robotic hands are essential for performing complex manipulation tasks, yet remain difficult to train due to the challenges of demonstration collection and high-dimensional control. While reinforcement learning (RL) can alleviate the data bottleneck by generating experience in simulation, it typically relies on carefully designed, task-specific reward functions, which hinder scalability and generalization. Thus, contemporary works in dexterous manipulation have often bootstrapped from reference trajectories. These trajectories specify target hand poses that guide the exploration of RL policies and object poses that enable dense, task-agnostic rewards. However, sourcing suitable trajectories-particularly for dexterous hands-remains a significant challenge. Yet, the precise details in explicit reference trajectories are often unnecessary, as RL ultimately refines the motion.
Our key insight is that modern vision-language models (VLMs) already encode the commonsense spatial and semantic knowledge needed to specify tasks and guide exploration effectively. Given a task description (e.g., "open the cabinet") and a visual scene, our method uses an off-the-shelf VLM to first identify task-relevant keypoints (e.g., handles, buttons) and then synthesize 3D trajectories for hand motion and object motion. Subsequently, we train a low-level residual RL policy in simulation to track these coarse trajectories or "scaffolds" with high fidelity. Across a number of simulated tasks involving articulated objects and semantic understanding, we demonstrate that our method is able to learn robust dexterous manipulation policies. Moreover, we showcase that our method transfers to realworld robotic hands without any human demonstrations or handcrafted rewards. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Vision-Language-Action Pretraining from Large-Scale Human VideosHao Luo, Yicheng Feng, Wanpeng Zhang, Sipeng Zheng 等ICML 2026 · 被引用 104 次
- GAE: Unleashing Physical Potential of VLM with Generalizable Action ExpertMingyu Liu, Zheng Huang, Xiaoyi Lin, Muzhi Zhu 等ICML 2026
它引用的顶会 Paper16
- Habitat 2.0: Training Home Assistants to Rearrange their HabitatAndrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans 等NeurIPS 2021 · 被引用 826 次
- Eureka: Human-Level Reward Design via Coding Large Language ModelsYecheng Jason Ma, William Liang, Guanzhi Wang, De-An Huang 等ICLR 2024 · 被引用 582 次
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng 等EMNLP 2024 · 被引用 479 次
- FoundationPose: Unified 6D Pose Estimation and Tracking of Novel ObjectsBowen Wen, Wei Yang, Jan Kautz, Stan BirchfieldCVPR 2024 · 被引用 215 次
- PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMsSoroush Nasiriany, Fei Xia, Wenhao Yu, Ted Xiao 等ICML 2024 · 被引用 212 次
相关 Paper
- Dexterous Manipulation Transfer via Progressive Kinematic-Dynamic AlignmentWenbin Bai, Qiyu Chen, Xiangbo Lin, Jianwen Li 等AAAI 2026
- Gentle Manipulation Policy Learning via Demonstrations from VLM Planned Atomic SkillsJiayu Zhou, Qiwei Wu, Jian Li, Zhe Chen 等AAAI 2026 · 被引用 1 次
- DemoGrasp: Universal Dexterous Grasping from a Single DemonstrationHaoqi Yuan, Ziye Huang, Ye Wang, Chuan Mao 等ICLR 2026 · 被引用 14 次
- SimpleVLA-RL: Scaling VLA Training via Reinforcement LearningHaozhan Li, Yuxin Zuo, Jiale Yu, Yuhao Zhang 等ICLR 2026 · 被引用 170 次
- ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language ModelsZirui Song, Guangxian Ouyang, Mingzhe Li, Yuheng Ji 等AAAI 2026 · 被引用 21 次
