IRASim: A Fine-Grained World Model for Robot Manipulation
Fangqi Zhu, Hongtao Wu, Song Guo, Yuxiao Liu, Chilam Cheang, Tao Kong
摘要
World models allow autonomous agents to plan and explore by predicting the visual outcomes of different actions. However, for robot manipulation, it is challenging to accurately model the fine-grained robot-object interaction within the visual space using existing methods which overlooks precise alignment between each action and the corresponding frame. In this paper, we present IRASim, a novel world model capable of generating videos with fine-grained robotobject interaction details, conditioned on historical observations and robot action trajectories. We train a diffusion transformer and introduce a novel frame-level actionconditioning module within each transformer block to explicitly model and strengthen the action-frame alignment. Extensive experiments show that: (1) the quality of the videos generated by our method surpasses all the baseline methods and scales effectively with increased model size and computation; (2) policy evaluations using IRASim exhibit a strong correlation with those using the ground-truth simulator, highlighting its potential to accelerate real-world policy evaluation; (3) testing-time scaling through model-based planning with IRASim significantly enhances policy performance, as evidenced by an improvement in the IoU metric on the Push-T benchmark from 0.637 to 0.961; (4) IRASim provides flexible action controllability, allowing virtual robotic arms in datasets to be controlled via a keyboard or VR controller. Video and code are available at https://gen-irasim.github.io/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- DreamDojo: A Real-Time Robot World Model from Large-Scale Human VideosShenyuan Gao, William Liang, Kaiyuan Zheng, Ayaan Malik 等ICML 2026 · 被引用 96 次
- WMPO: World Model-based Policy Optimization for Vision-Language-Action ModelsFangqi Zhu, Zhengyang Yan, Zicong Hong, Quanxin Shou 等ICLR 2026 · 被引用 64 次
- RoboScape: Physics-informed Embodied World ModelYu Shang, Xin Zhang, Yinzhou Tang, Lei Jin 等NeurIPS 2025 · 被引用 43 次
- World Guidance: World Modeling in Condition Space for Action GenerationYue Su, Sijin Chen, Haixin Shi, Mingyu Liu 等ICML 2026 · 被引用 26 次
- ORV: 4D Occupancy-centric Robot Video GenerationXiuyu Yang, Bohan Li, Shaocong Xu, Nan Wang 等CVPR 2026 · 被引用 19 次
它引用的顶会 Paper18
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- Pseudo Numerical Methods for Diffusion Models on ManifoldsLuping Liu, Yi Ren, Zhijie Lin, Zhou ZhaoICLR 2022 · 被引用 861 次
相关 Paper
- Pre-Trained Video Generative Models as World SimulatorsHaoran He, Yang Zhang, Liang Lin, Zhongwen Xu 等AAAI 2026 · 被引用 32 次
- WorldGym: World Model as An Environment for Policy EvaluationJulian Quevedo, Ansh Kumar Sharma, Yixiang Sun, Varad Suryavanshi 等ICLR 2026 · 被引用 58 次
- GWM: Towards Scalable Gaussian World Models for Robotic ManipulationGuanxing Lu, Baoxiong Jia, Puhao Li, Yixin Chen 等ICCV 2025 · 被引用 1 次
- Ctrl-World: A Controllable Generative World Model for Robot ManipulationYanjiang Guo, Lucy Xiaoyang Shi, Jianyu Chen, Chelsea FinnICLR 2026 · 被引用 163 次
- ReSim: Reliable World Simulation for Autonomous DrivingJiazhi Yang, Kashyap Chitta, Shenyuan Gao, Long Chen 等NeurIPS 2025 · 被引用 53 次
