H-InDex: Visual Reinforcement Learning with Hand-Informed Representations for Dexterous Manipulation
Yanjie Ze, Yuyao Liu, Ruizhe Shi, Jiaxin Qin, Zhecheng Yuan, Jiashun Wang, Huazhe Xu
Abstract
Human hands possess remarkable dexterity and have long served as a source of inspiration for robotic manipulation. In this work, we propose a human Hand-Informed visual representation learning framework to solve difficult Dexterous manipulation tasks (H-InDex) with reinforcement learning. Our framework consists of three stages: (i) pre-training representations with 3D human hand pose estimation, (ii) offline adapting representations with self-supervised keypoint detection, and (iii) reinforcement learning with exponential moving average BatchNorm. The last two stages only modify 0.36% parameters of the pre-trained representation in total, ensuring the knowledge from pre-training is maintained to the full extent. We empirically study 12 challenging dexterous manipulation tasks and find that H-InDex largely surpasses strong baseline methods and the recent visual foundation models for motor control. Code is available at yanjieze.com/H-InDex. * Model params: ResNet-50 (24 M), ViT-B (86 M), ViT-S (22 M) Normalized Average Score H-InDex (ResNet-50) RRL (ResNet-50) R3M (ResNet-50) VC-1 (ViT-B) MVP (ViT-S)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 325d9c5e-3a48-427c-8f96-3ee2a3a19ecdCited by top-tier papers5
- Unleashing the Power of Pre-trained Language Models for Offline Reinforcement LearningRuizhe Shi, Yuyao Liu, Yanjie Ze, Simon Shaolei Du et al.ICLR 2024 · 36 citations
- Uni-O4: Unifying Online and Offline Deep Reinforcement Learning with Multi-Step On-Policy OptimizationKun Lei, Zhengmao He, Chenhao Lu, Kaizhe Hu et al.ICLR 2024 · 31 citations
- When Should We Prefer State-to-Visual DAgger over Visual Reinforcement Learning?Tongzhou Mu, Zhaoyang Li, Stanislaw Wiktor Strzelecki, Xiu Yuan et al.AAAI 2025 · 8 citations
- Causal Information Prioritization for Efficient Reinforcement LearningHongye Cao, Fan Feng, Tianpei Yang, Jing Huo et al.ICLR 2025 · 1 citation
- RAPID Hand: Robust, Affordable, Perception-Integrated, Dexterous Manipulation Platform for Embodied IntelligenceZhaoliang Wan, Zetong Bi, Zida Zhou, Hao Ren et al.NeurIPS 2025 · 1 citation
Builds on9
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?Arjun Majumdar, Karmesh Yadav, Sergio Arnaud, Yecheng Jason Ma et al.NeurIPS 2023 · 336 citations
- The Unsurprising Effectiveness of Pre-Trained Vision Models for ControlSimone Parisi, Aravind Rajeswaran, Senthil Purushwalkam, Abhinav GuptaICML 2022 · 233 citations
- RRL: Resnet as representation for Reinforcement LearningRutav M. Shah, Vikash KumarICML 2021 · 129 citations
- On Pre-Training for Visuo-Motor Control: Revisiting a Learning-from-Scratch BaselineNicklas Hansen, Zhecheng Yuan, Yanjie Ze, Tongzhou Mu et al.ICML 2023 · 78 citations
Related papers
- VRL3: A Data-Driven Framework for Visual Deep Reinforcement LearningChe Wang, Xufang Luo, Keith W. Ross, Dongsheng LiNeurIPS 2022 · 72 citations
- Curious Representation Learning for Embodied IntelligenceYilun Du, Chuang Gan, Phillip IsolaICCV 2021 · 50 citations
- VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-TrainingYecheng Jason Ma, Shagun Sodhani, Dinesh Jayaraman, Osbert Bastani et al.ICLR 2023 · 35 citations
- Unsupervised Learning of Visual 3D Keypoints for ControlBoyuan Chen, Pieter Abbeel, Deepak PathakICML 2021 · 46 citations
- HDP: Triply‑Hierarchical Diffusion Policy for Visuomotor LearningYiyang Lu, Yufeng Tian, Zhecheng Yuan, Xianbang Wang et al.ICLR 2026 · 10 citations
