VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception
Zhaoliang Wan, Yonggen Ling, Senlin Yi, Lu Qi, Wang Wei Lee, Minglei Lu, Sicheng Yang, Xiao Teng, Peng Lu, Xu Yang, Ming-Hsuan Yang, Hui Cheng
Abstract
This paper addresses the scarcity of large-scale datasets for accurate object-in-hand pose estimation, which is crucial for robotic in-hand manipulation within the ``Perception-Planning-Control"paradigm. Specifically, we introduce VinT-6D, the first extensive multi-modal dataset integrating vision, touch, and proprioception, to enhance robotic manipulation. VinT-6D comprises 2 million VinT-Sim and 0.1 million VinT-Real splits, collected via simulations in MuJoCo and Blender and a custom-designed real-world platform. This dataset is tailored for robotic hands, offering models with whole-hand tactile perception and high-quality, well-aligned data. To the best of our knowledge, the VinT-Real is the largest considering the collection difficulties in the real-world environment so that it can bridge the gap of simulation to real compared to the previous works. Built upon VinT-6D, we present a benchmark method that shows significant improvements in performance by fusing multi-modal information. The project is available at https://VinT-6D.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 890a020e-57be-43de-b886-960a565b81a9Cited by top-tier papers4
- Conditional Panoramic Image Generation via Masked Autoregressive ModelingChaoyang Wang, Xiangtai Li, Lu Qi, Xiaofan Lin et al.NeurIPS 2025 · 11 citations
- RAPID Hand: Robust, Affordable, Perception-Integrated, Dexterous Manipulation Platform for Embodied IntelligenceZhaoliang Wan, Zetong Bi, Zida Zhou, Hao Ren et al.NeurIPS 2025 · 1 citation
- Cross-Tactile Sensor Representation LearningYan Zhang, Zheng WANG, Pengpeng Zeng, Xing Xu et al.ICML 2026
- Layout-your-3D: Controllable and Precise 3D Generation with 2D BlueprintJunwei Zhou, Xueting Li, Lu Qi, Ming-Hsuan YangICLR 2025
Builds on9
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- CDPN: Coordinates-Based Disentangled Pose Network for Real-Time RGB-Based 6-DoF Object Pose EstimationZhigang Li, Gu Wang, Xiangyang JiICCV 2019 · 482 citations
- EPro-PnP: Generalized End-to-End Probabilistic Perspective-n-Points for Monocular Object Pose EstimationHansheng Chen, Pichao Wang, Fan Wang, Wei Tian et al.CVPR 2022 · 175 citations
- UniDexGrasp++: Improving Dexterous Grasping Policy Learning via Geometry-aware Curriculum and Iterative Generalist-Specialist LearningWeikang Wan, Haoran Geng, Yun Liu, Zikang Shan et al.ICCV 2023 · 160 citations
- High Quality Entity SegmentationLu Qi, Jason Kuen, Tiancheng Shen, Jiuxiang Gu et al.ICCV 2023 · 91 citations
Related papers
- VTDexManip: A Dataset and Benchmark for Visual-tactile Pretraining and Dexterous Manipulation with Reinforcement LearningQingtao Liu, Yu Cui, Zhengnan Sun, Gaofeng Li et al.ICLR 2025
- Visual-Tactile Sensing for In-Hand Object ReconstructionWenqiang Xu, Zhenjun Yu, Han Xue, Ruolin Ye et al.CVPR 2023
- 3D Shape Reconstruction from Vision and TouchEdward J. Smith, Roberto Calandra, Adriana Romero, Georgia Gkioxari et al.NeurIPS 2020 · 90 citations
- DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile AdapterXukun Li, Yu Sun, Lei Zhang, Bo-Sheng Huang et al.ICML 2026
- AnyTouch 2: General Optical Tactile Representation Learning For Dynamic Tactile PerceptionRuoxuan Feng, Yuxuan Zhou, Siyu Mei, Dongzhan Zhou et al.ICLR 2026 · 25 citations
