GenSim: Generating Robotic Simulation Tasks via Large Language Models
Lirui Wang, Yiyang Ling, Zhecheng Yuan, Mohit Shridhar, Chen Bao, Yuzhe Qin, Bailin Wang, Huazhe Xu, Xiaolong Wang
Abstract
Collecting large amounts of real-world interaction data to train general robotic policies is often prohibitively expensive, thus motivating the use of simulation data. However, existing methods for data generation have generally focused on scene-level diversity (e.g., object instances and poses) rather than task-level diversity, due to the human effort required to come up with and verify novel tasks. This has made it challenging for policies trained on simulation data to demonstrate significant task-level generalization. In this paper, we propose to automatically generate rich simulation environments and expert demonstrations by exploiting a large language models' (LLM) grounding and coding ability. Our approach, dubbed GenSim, has two modes: goal-directed generation, wherein a target task is given to the LLM and the LLM proposes a task curriculum to solve the target task, and exploratory generation, wherein the LLM bootstraps from previous tasks and iteratively proposes novel tasks that would be helpful in solving more complex tasks. We use GPT4 to expand the existing benchmark by ten times to over 100 tasks, on which we conduct supervised finetuning and evaluate several LLMs including finetuned GPTs and Code Llama on code generation for robotic simulation tasks. Furthermore, we observe that LLMs-generated simulation programs can enhance task-level generalization significantly when used for multitask policy training. We further find that with minimal sim-to-real adaptation, the multitask policies pretrained on GPT4-generated simulation tasks exhibit stronger transfer to unseen long-horizon tasks in the real world and outperform baselines by 25%. See the project website (https://liruiw.github.io/gensim) for code, demos, and videos.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0974b693-cebc-431f-b106-1b951a2fd7fdCited by top-tier papers27
- ReEvo: Large Language Models as Hyper-Heuristics with Reflective EvolutionHaoran Ye, Jiarui Wang, Zhiguang Cao, Federico Berto et al.NeurIPS 2024 · 424 citations
- Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained TransformersLirui Wang, Xinlei Chen, Jialiang Zhao, Kaiming HeNeurIPS 2024 · 208 citations
- RoboCasa365: A Large-Scale Simulation Framework for Training and Benchmarking Generalist RobotsSoroush Nasiriany, Sepehr Nasiriany, Abhiram Maddukuri, Yuke ZhuICLR 2026 · 96 citations
- Closed-Loop Visuomotor Control with Generative Expectation for Robotic ManipulationQingwen Bu, Jia Zeng, Li Chen, Yanchao Yang et al.NeurIPS 2024 · 80 citations
- Architect: Generating Vivid and Interactive 3D Scenes with Hierarchical 2D InpaintingYian Wang, Xiaowen Qiu, Jiageng Liu, Zhehuan Chen et al.NeurIPS 2024 · 48 citations
Builds on5
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- 🏘️ ProcTHOR: Large-Scale Embodied AI Using Procedural GenerationMatt Deitke, Eli VanderBilt, Alvaro Herrasti, Luca Weihs et al.NeurIPS 2022 · 596 citations
- DreamFusion: Text-to-3D using 2D DiffusionBen Poole, Ajay Jain, Jonathan T. Barron, Ben MildenhallICLR 2023 · 463 citations
- Mind's Eye: Grounded Language Model Reasoning through SimulationRuibo Liu, Jason Wei, Shixiang Shane Gu, Te-Yen Wu et al.ICLR 2023 · 22 citations
Related papers
- GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following ManipulationNing Gao, Yilun Chen, Shuai Yang, Xinyi Chen et al.CVPR 2025
- TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task TypesJiankang Chen, Tianke Zhang, Changyi Liu, Haojie Ding et al.ICLR 2025
- Program Synthesis Benchmark for Visual Programming in XLogoOnline EnvironmentChao Wen, Jacqueline Staub, Adish SinglaACL 2025 · 6 citations
- VIMA: Robot Manipulation with Multimodal PromptsYunfan Jiang, Agrim Gupta, Zichen Zhang, Guanzhi Wang et al.ICML 2023 · 80 citations
- VirtualEnv: A Platform for Embodied AI ResearchKabir Swain, Sijie Han, Ayush Raina, Jin Zhang et al.AAAI 2026
