OakInk2 : A Dataset of Bimanual Hands-Object Manipulation in Complex Task Completion
Xinyu Zhan, Lixin Yang, Yifei Zhao, Kangrui Mao, Hanlin Xu, Zenan Lin, Kailin Li, Cewu Lu
2024年份
31顶会引用
摘要
Use the knife to cut the apple; then use the clamp to grip the sugar cubes into the bowl; afterwards, use the microwave oven to heat the bowl.
L2 L3 L9 OPEN CLOSE OPEN CLOSE CLOSE OPEN <heat, sth> Figure 1. An overview of the data and content of our proposed OAKINK2 dataset. OAKINK2 dataset focuses on bimanual object manipulation tasks for complex daily activities. 1) The top row shows the data collection process, including the task setup (top-left panel), human demonstration (top-center), and annotation (top-right).
- The second row shows the three levels of abstraction constructed by OAKINK2 for complex tasks, including the Affordance, Primitive Task, and Complex Task. OAKINK2 dataset provides allocentric and egocentric videos of human manipulation process, as well as the corresponding 3D-pose annotation and task specification.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- Vision-Language-Action Pretraining from Large-Scale Human VideosHao Luo, Yicheng Feng, Wanpeng Zhang, Sipeng Zheng 等ICML 2026 · 被引用 104 次
- OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language ModelZhenhao Zhang, Ye Shi, Lingxiao Yang, Suting Ni 等NeurIPS 2025 · 被引用 25 次
- CoDA: Coordinated Diffusion Noise Optimization for Whole-Body Manipulation of Articulated ObjectsHuaijin Pi, Zhi Cen, Zhiyang Dou, Taku KomuraNeurIPS 2025 · 被引用 14 次
- Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human VideosYicheng Feng, Wanpeng Zhang, Ye Wang, Hao Luo 等CVPR 2026 · 被引用 14 次
- MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion GenerationBohan Zhou, Yi Zhan, Zhongbin Zhang, Zongqing LuNeurIPS 2025 · 被引用 14 次
它引用的顶会 Paper33
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll 等ICCV 2019 · 被引用 1,784 次
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 被引用 1,539 次
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 被引用 672 次
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis 等CVPR 2022 · 被引用 525 次
- FreiHAND: A Dataset for Markerless Capture of Hand Pose and Shape From Single RGB ImagesChristian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan C. Russell 等ICCV 2019 · 被引用 493 次
相关 Paper
- OakInk: A Large-scale Knowledge Repository for Understanding Hand-Object InteractionLixin Yang, Kailin Li, Xinyu Zhan, Fei Wu 等CVPR 2022 · 被引用 79 次
- TACO: Benchmarking Generalizable Bimanual Tool-ACtion-Object UnderstandingYun Liu, Haolin Yang, Xu Si, Ling Liu 等CVPR 2024 · 被引用 12 次
- EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric VideoRyan Hoque, Peide Huang, David J. Yoon, Mouli Sivapurapu 等ICLR 2026 · 被引用 248 次
- ARCTIC: A Dataset for Dexterous Bimanual Hand-Object ManipulationZicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas 等CVPR 2023
- Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric VideosChiara Plizzari, Alessio Tonioni, Yongqin Xian, Achin Kulshrestha 等CVPR 2025
