GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation
Ning Gao, Yilun Chen, Shuai Yang, Xinyi Chen, Yang Tian, Hao Li, Haifeng Huang, Hanqing Wang, Tai Wang, Jiangmiao Pang
Abstract
Robotic manipulation in real-world settings remains challenging, especially regarding robust generalization. Existing simulation platforms lack sufficient support for exploring how policies adapt to varied instructions and scenarios. Thus, they lag behind the growing interest in instructionfollowing foundation models like LLMs, whose adaptability is crucial yet remains underexplored in fair comparisons. To bridge this gap, we introduce GENMANIP, a realistic tabletop simulation platform tailored for policy generalization studies. It features an automatic pipeline via LLMdriven task-oriented scene graph to synthesize large-scale, diverse tasks using 10K annotated 3D object assets. To systematically assess generalization, we present GENMANIP-BENCH, a benchmark of 200 scenarios refined via humanin-the-loop corrections. We evaluate two policy types: (1) modular manipulation systems integrating foundation models for perception, reasoning, and planning, and (2) endto-end policies trained through scalable data collection. Results show that while data scaling benefits end-to-end methods, modular systems enhanced with foundation models generalize more effectively across diverse scenarios. We anticipate this platform to facilitate critical insights for advancing policy generalization in realistic conditions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 11b49128-5d9b-42d9-bd8e-0a9bf99ef4edCited by top-tier papers7
- Vision-Language-Action Instruction Tuning: From Understanding to ManipulationShuai Yang, Hao Li, Bin Wang, Yilun Chen et al.ICLR 2026 · 50 citations
- RoboInter: A Holistic Intermediate Representation Suite Towards Robotic ManipulationHao Li, Ziqin Wang, Zi-han Ding, Shuai Yang et al.ICLR 2026 · 17 citations
- MM-ACT: Learn from Multimodal Parallel Generation to ActHaotian Liang, Xinyi Chen, Bin Wang, Mingkang Chen et al.CVPR 2026 · 13 citations
- Spatially Guided Training for Vision-Language-Action ModelJinhui Ye, Fangjing Wang, Ning Gao, Junqiu Yu et al.ICLR 2026 · 6 citations
- STABLE: Simulation-Ready Tabletop Layout Generation via a Semantics–Physics Dual SystemZhen Luo, Yixuan Yang, Xudong XU, Jinkun Hao et al.ICML 2026
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov et al.ICCV 2023 · 1,662 citations
- 🏘️ ProcTHOR: Large-Scale Embodied AI Using Procedural GenerationMatt Deitke, Eli VanderBilt, Alvaro Herrasti, Luca Weihs et al.NeurIPS 2022 · 596 citations
- Vision-Language Foundation Models as Effective Robot ImitatorsXinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu et al.ICLR 2024 · 375 citations
- An Embodied Generalist Agent in 3D WorldJiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu et al.ICML 2024 · 361 citations
Related papers
- Genie Envisioner: A Unified World Foundation Platform for Robotic ManipulationYue Liao, Pengfei Zhou, Siyuan Huang, Donglin Yang et al.ICLR 2026 · 136 citations
- GenSim: Generating Robotic Simulation Tasks via Large Language ModelsLirui Wang, Yiyang Ling, Zhecheng Yuan, Mohit Shridhar et al.ICLR 2024 · 143 citations
- RoboCasa365: A Large-Scale Simulation Framework for Training and Benchmarking Generalist RobotsSoroush Nasiriany, Sepehr Nasiriany, Abhiram Maddukuri, Yuke ZhuICLR 2026 · 96 citations
- VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning TasksShiduo Zhang, Zhe Xu, Peiju Liu, Xiaopeng Yu et al.ICCV 2025 · 12 citations
- SimWorld-Robotics: Synthesizing Photorealistic and Dynamic Urban Environments for Multimodal Robot Navigation and CollaborationYan Zhuang, Jiawei Ren, Xiaokang Ye, Jianzhi Shen et al.NeurIPS 2025 · 9 citations
