OakInk2 : A Dataset of Bimanual Hands-Object Manipulation in Complex Task Completion
Xinyu Zhan, Lixin Yang, Yifei Zhao, Kangrui Mao, Hanlin Xu, Zenan Lin, Kailin Li, Cewu Lu
2024Year
31Top-tier citations
Abstract
Use the knife to cut the apple; then use the clamp to grip the sugar cubes into the bowl; afterwards, use the microwave oven to heat the bowl.
L2 L3 L9 OPEN CLOSE OPEN CLOSE CLOSE OPEN <heat, sth> Figure 1. An overview of the data and content of our proposed OAKINK2 dataset. OAKINK2 dataset focuses on bimanual object manipulation tasks for complex daily activities. 1) The top row shows the data collection process, including the task setup (top-left panel), human demonstration (top-center), and annotation (top-right).
- The second row shows the three levels of abstraction constructed by OAKINK2 for complex tasks, including the Affordance, Primitive Task, and Complex Task. OAKINK2 dataset provides allocentric and egocentric videos of human manipulation process, as well as the corresponding 3D-pose annotation and task specification.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers31
- Vision-Language-Action Pretraining from Large-Scale Human VideosHao Luo, Yicheng Feng, Wanpeng Zhang, Sipeng Zheng et al.ICML 2026 · 104 citations
- OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language ModelZhenhao Zhang, Ye Shi, Lingxiao Yang, Suting Ni et al.NeurIPS 2025 · 25 citations
- CoDA: Coordinated Diffusion Noise Optimization for Whole-Body Manipulation of Articulated ObjectsHuaijin Pi, Zhi Cen, Zhiyang Dou, Taku KomuraNeurIPS 2025 · 14 citations
- Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human VideosYicheng Feng, Wanpeng Zhang, Ye Wang, Hao Luo et al.CVPR 2026 · 14 citations
- MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion GenerationBohan Zhou, Yi Zhan, Zhongbin Zhang, Zongqing LuNeurIPS 2025 · 14 citations
Builds on33
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 672 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- FreiHAND: A Dataset for Markerless Capture of Hand Pose and Shape From Single RGB ImagesChristian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan C. Russell et al.ICCV 2019 · 493 citations
Related papers
- OakInk: A Large-scale Knowledge Repository for Understanding Hand-Object InteractionLixin Yang, Kailin Li, Xinyu Zhan, Fei Wu et al.CVPR 2022 · 79 citations
- TACO: Benchmarking Generalizable Bimanual Tool-ACtion-Object UnderstandingYun Liu, Haolin Yang, Xu Si, Ling Liu et al.CVPR 2024 · 12 citations
- EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric VideoRyan Hoque, Peide Huang, David J. Yoon, Mouli Sivapurapu et al.ICLR 2026 · 248 citations
- ARCTIC: A Dataset for Dexterous Bimanual Hand-Object ManipulationZicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas et al.CVPR 2023
- Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric VideosChiara Plizzari, Alessio Tonioni, Yongqin Xian, Achin Kulshrestha et al.CVPR 2025
