UniHM: Unified Dexterous Hand Manipulation with Vision Language Model
Zhenhao Zhang, Jiaxin Liu, Ye Shi, Jingya Wang
摘要
Planning physically feasible dexterous hand manipulation is a central challenge in robotic manipulation and Embodied AI. Prior work typically relies on object-centric cues or precise hand-object interaction sequences, foregoing the rich, compositional guidance of open-vocabulary instruction. We introduce UniHM, the first framework for unified dexterous hand manipulation guided by free-form language commands. We propose a Unified Hand-Dexterous Tokenizer that maps heterogeneous dexterous-hand morphologies into a single shared codebook, improving cross-dexterous hand generalization and scalability to new morphologies. Our vision language action model is trained solely on human-object interaction data, eliminating the need for massive real-world teleoperation datasets, and demonstrates strong generalizability in producing human-like manipulation sequences from open-ended language instructions. To ensure physical realism, we introduce a physics-guided dynamic refinement module that performs segment-wise joint optimization under generative and temporal priors, yielding smooth and physically feasible manipulation sequences. Across multiple datasets and real-world evaluations, UniHM attains state-of-the-art results on both seen and unseen objects and trajectories, demonstrating strong generalization and high physical feasibility. Our project page at https://unihm.github.io/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper18
- SnowflakeNet: Point Cloud Completion by Snowflake Point Deconvolution with Skip-TransformerPeng Xiang, Xin Wen, Yu-Shen Liu, Yan-Pei Cao 等ICCV 2021 · 被引用 318 次
- Grasp as You Say: Language-guided Dexterous Grasp GenerationYi-Lin Wei, Jian-Jian Jiang, Chengyi Xing, Xiantuo Tan 等NeurIPS 2024 · 被引用 85 次
- StreamForest: Efficient Online Video Understanding with Persistent Event MemoryXiangyu Zeng, Kefan Qiu, Qingyu Zhang, Xinhao Li 等NeurIPS 2025 · 被引用 79 次
- OakInk: A Large-scale Knowledge Repository for Understanding Hand-Object InteractionLixin Yang, Kailin Li, Xinyu Zhan, Fei Wu 等CVPR 2022 · 被引用 79 次
- MotionGPT3: Human Motion as a Second ModalityBingfan Zhu, Biao Jiang, Sunyi Wang, Shixiang Tang 等ICLR 2026 · 被引用 43 次
相关 Paper
- Vision-Language-Action Pretraining from Large-Scale Human VideosHao Luo, Yicheng Feng, Wanpeng Zhang, Sipeng Zheng 等ICML 2026 · 被引用 104 次
- OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language ModelZhenhao Zhang, Ye Shi, Lingxiao Yang, Suting Ni 等NeurIPS 2025 · 被引用 25 次
- Unified Human-Scene Interaction via Prompted Chain-of-ContactsZeqi Xiao, Tai Wang, Jingbo Wang, Jinkun Cao 等ICLR 2024 · 被引用 113 次
- Cross-Hand Latent Representation for Vision-Language-Action ModelsGuangqi Jiang, Yutong Liang, Jianglong Ye, Jia-Yang Huang 等CVPR 2026 · 被引用 14 次
- UniDex: A Robot Foundation Suite for Universal Dexterous Hand Control from Egocentric Human VideosGu Zhang, Qicheng Xu, Haozhe Zhang, Jianhan Ma 等CVPR 2026 · 被引用 23 次
