VTDexManip: A Dataset and Benchmark for Visual-tactile Pretraining and Dexterous Manipulation with Reinforcement Learning
Qingtao Liu, Yu Cui, Zhengnan Sun, Gaofeng Li, Jiming Chen, Qi Ye
Abstract
Vision and touch are the most commonly used senses in human manipulation. While leveraging human manipulation videos for robotic task pretraining has shown promise in prior works, it is limited to image and language modalities and deployment to simple parallel grippers. In this paper, aiming to address the limitations, we collect a vision-tactile dataset by humans manipulating 10 daily tasks and 182 objects. In contrast with the existing datasets, our dataset is the first visual-tactile dataset for complex robotic manipulation skill learning. Also, we introduce a novel benchmark, featuring six complex dexterous manipulation tasks and a reinforcement learning-based vision-tactile skill learning framework. 18 non-pretraining and pretraining methods within the framework are designed and compared to investigate the effectiveness of different modalities and pertaining strategies. Key findings based on our benchmark results and analyses experiments include: 1) Despite the tactile modality used in our experiments being binary and sparse, including it directly in the policy training boosts the success rate by about 20% and joint pretraining it with vision gains a further 20%. 2) Joint pretraining visual-tactile modalities exhibits strong adaptability in unknown tasks and achieves robust performance among all tasks. 3) Using binary tactile signals with vision is robust to viewpoint setting, tactile noise, and the binarization threshold, which facilitates to the visual-tactile policy to be deployed in reality. The dataset and benchmark are available at https://github.com/LQTS/VTDexManip . * Coresponing Author • We collect a human visual-tactile manipulation dataset consisting of 565k frames, covering 10 daily tasks and 182 objects for multi-fingered robotic hand manipulation. • We propose a vision-tactile benchmark for dexterous manipulation, which includes a manipulation simulation platform with six multi-fingered manipulation tasks and a manipulation skill learning framework based on pretraining and reinforcement learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b8fa1556-df41-4e12-928a-7fa4000a3a5dCited by top-tier papers3
- TaCo: A Benchmark for Lossless and Lossy Codecs of Heterogeneous Tactile DataZhengxue Cheng, Yan Zhao, Keyu Wang, Hengdi Zhang et al.ICLR 2026 · 3 citations
- EgoTactile: Learning Grasp Pressure for Everyday Objects from Egocentric VideoYuan Zeng, Yujia Shi, Tiao Tan, Xingting Li et al.ICML 2026 · 1 citation
- DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile AdapterXukun Li, Yu Sun, Lei Zhang, Bo-Sheng Huang et al.ICML 2026
Builds on8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- LIV: Language-Image Representations and Rewards for Robotic ControlYecheng Jason Ma, Vikash Kumar, Amy Zhang, Osbert Bastani et al.ICML 2023 · 212 citations
- VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-TrainingYecheng Jason Ma, Shagun Sodhani, Dinesh Jayaraman, Osbert Bastani et al.ICLR 2023 · 35 citations
Related papers
- VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and ProprioceptionZhaoliang Wan, Yonggen Ling, Senlin Yi, Lu Qi et al.ICML 2024 · 11 citations
- DexMove: Learning Tactile-Guided Non-Prehensile Manipulation with Dexterous HandsPei Lin, Yuzhe Huang, Wanlin Li, Chenxi Xiao et al.ICLR 2026
- Visual-Tactile Sensing for In-Hand Object ReconstructionWenqiang Xu, Zhenjun Yu, Han Xue, Ruolin Ye et al.CVPR 2023
- Benchmarking Offline Reinforcement Learning on Real-Robot HardwareNico Gürtler, Sebastian Blaes, Pavel Kolev, Felix Widmaier et al.ICLR 2023 · 11 citations
- Universal Visuo-Tactile Video Understanding for Embodied InteractionYifan Xie, Mingyang Li, Shoujie Li, Xingting Li et al.NeurIPS 2025 · 16 citations
