ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills
Jiayuan Gu, Fanbo Xiang, Xuanlin Li, Zhan Ling, Xiqiang Liu, Tongzhou Mu, Yihe Tang, Stone Tao, Xinyue Wei, Yunchao Yao, Xiaodi Yuan, Pengwei Xie
Abstract
Generalizable manipulation skills, which can be composed to tackle longhorizon and complex daily chores, are one of the cornerstones of Embodied AI. However, existing benchmarks, mostly composed of a suite of simulatable environments, are insufficient to push cutting-edge research works because they lack object-level topological and geometric variations, are not based on fully dynamic simulation, or are short of native support for multiple types of manipulation tasks. To this end, we present ManiSkill2, the next generation of the SAPIEN ManiSkill benchmark, to address critical pain points often encountered by researchers when using benchmarks for generalizable manipulation skills. ManiSkill2 includes 20 manipulation task families with 2000+ object models and 4M+ demonstration frames, which cover stationary/mobile-base, single/dualarm, and rigid/soft-body manipulation tasks with 2D/3D-input data simulated by fully dynamic engines. It defines a unified interface and evaluation protocol to support a wide range of algorithms (e.g., classic sense-plan-act, RL, IL), visual observations (point cloud, RGBD), and controllers (e.g., action type and parameterization). Moreover, it empowers fast visual input learning algorithms so that a CNN-based policy can collect samples at about 2000 FPS with 1 GPU and 16 processes on a regular workstation. It implements a render server infrastructure to allow sharing rendering resources across all environments, thereby significantly reducing memory usage. We open-source all codes of our benchmark (simulator, environments, and baselines) and host an online challenge open to interdisciplinary researchers. Figure 1 : ManiSkill2 provides a unified, fast, and accessible system that encompasses well-curated manipulation tasks (e.g., stationary/mobile-base, single/dual-arm, rigid/soft-body). 1 † and * indicate equal contribution (in alphabetical order). See Appendix H for contributions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ade179df-b035-435e-a362-8eddaf1cccddCited by top-tier papers106
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationTianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai et al.ICML 2026 · 394 citations
- TD-MPC2: Scalable, Robust World Models for Continuous ControlNicklas Hansen, Hao Su, Xiaolong WangICLR 2024 · 388 citations
- RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative SimulationYufei Wang, Zhou Xian, Feng Chen, Tsun-Hsuan Wang et al.ICML 2024 · 227 citations
- HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action ModelJiaming Liu, Hao Chen, Zhuoyang Liu, Pengju An et al.ICLR 2026 · 216 citations
- Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action ModelFuhao Li, Wenxuan Song, Han Zhao, Jingbo Wang et al.ICLR 2026 · 145 citations
Builds on6
- Habitat 2.0: Training Home Assistants to Rearrange their HabitatAndrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans et al.NeurIPS 2021 · 826 citations
- Approximate convex decomposition for 3D meshes with collision-aware concavity and tree searchXinyue Wei, Minghua Liu, Zhan Ling, Hao SuSIGGRAPH 2022 · 79 citations
- Improving Policy Optimization with Generalist-Specialist LearningZhiwei Jia, Xuanlin Li, Zhan Ling, Shuang Liu et al.ICML 2022 · 32 citations
- Multi-skill Mobile Manipulation for Object RearrangementJiayuan Gu, Devendra Singh Chaplot, Hao Su, Jitendra MalikICLR 2023 · 10 citations
- ManipulaTHOR: A Framework for Visual Object ManipulationKiana Ehsani, Winson Han, Alvaro Herrasti, Eli VanderBilt et al.CVPR 2021
Related papers
- Rapid Motor Adaptation for Robotic Manipulator ArmsYichao Liang, Kevin Ellis, João F. HenriquesCVPR 2024
- VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning TasksShiduo Zhang, Zhe Xu, Peiju Liu, Xiaopeng Yu et al.ICCV 2025 · 12 citations
- GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following ManipulationNing Gao, Yilun Chen, Shuai Yang, Xinyi Chen et al.CVPR 2025
- EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied AgentsRui Yang, Hanyang Chen, Junyu Zhang, Mark Zhao et al.ICML 2025
- SAGE: Scalable Agentic 3D Scene Generation for Embodied AIHongchi Xia, Xuan Li, Zhaoshuo Li, Qianli Ma et al.CVPR 2026 · 50 citations
