WorldGym: World Model as An Environment for Policy Evaluation
Julian Quevedo, Ansh Kumar Sharma, Yixiang Sun, Varad Suryavanshi, Percy Liang, Sherry Yang
摘要
Evaluating robot control policies is difficult: real-world testing is costly, and handcrafted simulators require manual effort to improve in realism and generality. We propose a world-model-based policy evaluation environment (WorldGym), an autoregressive, action-conditioned video generation model which serves as a proxy to real world environments. Policies are evaluated via Monte Carlo rollouts in the world model, with a vision-language model providing rewards. We evaluate a set of VLA-based real-robot policies in the world model using only initial frames from real robots, and show that policy success rates within the world model highly correlate with real-world success rates. Moreoever, we show that WorldGym is able to preserve relative policy rankings across different policy versions, sizes, and training checkpoints. Due to requiring only a single start frame as input, the world model further enables efficient evaluation of robot policies' generalization ability on novel tasks and environments. We find that modern VLA-based robot policies still struggle to distinguish object shapes and can become distracted by adversarial facades of objects. While generating highly realistic object interaction remains challenging, WorldGym faithfully emulates robot motions and offers a practical starting point for safe and reproducible policy evaluation before deployment. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- PointWorld: Scaling 3D World Models for In-The-Wild Robotic ManipulationWenlong Huang, Yu-Wei Chao, Arsalan Mousavian, Ming-Yu Liu 等CVPR 2026 · 被引用 87 次
- DreamSAC: Learning Hamiltonian World Models via Symmetry ExplorationJinzhou Tang, Fan Feng, Minghao Fu, Wenjun Lin 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper20
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 被引用 1,170 次
相关 Paper
- Scaling Real-World Robot Policy Evaluation via Discrete Diffusion World ModelYaxuan Li, Junjie Wen, Zhongyi Zhou, Yefei Chen 等ICML 2026 · 被引用 5 次
- IRASim: A Fine-Grained World Model for Robot ManipulationFangqi Zhu, Hongtao Wu, Song Guo, Yuxiao Liu 等ICCV 2025 · 被引用 4 次
- Ctrl-World: A Controllable Generative World Model for Robot ManipulationYanjiang Guo, Lucy Xiaoyang Shi, Jianyu Chen, Chelsea FinnICLR 2026 · 被引用 163 次
- VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World ModelYanjiang Guo, Tony Lee, Lucy Xiaoyang Shi, Jianyu Chen 等ICML 2026 · 被引用 29 次
- WMPO: World Model-based Policy Optimization for Vision-Language-Action ModelsFangqi Zhu, Zhengyang Yan, Zicong Hong, Quanxin Shou 等ICLR 2026 · 被引用 64 次
