Scaling Real-World Robot Policy Evaluation via Discrete Diffusion World Model
Yaxuan Li, Junjie Wen, Zhongyi Zhou, Yefei Chen, Chaomin Shen, Yaxin Peng, Yichen Zhu
Abstract
Evaluating generalist robot manipulation policies is costly and difficult to scale in the real world. While emerging world models (e.g., WorldEval, Ctrl-World) offer a promising alternative, the reliability of such evaluation remains a critical bottleneck. Specifically, their visual predictions can undermine policy assessment by "self-correcting" failures into false positives or yielding artifacts under out-of-distribution controls. Even with failure-enriched data, current architectures struggle to capture action-causal dynamics, as they typically treat actions as passive conditions rather than causal drivers. To address this, we propose dWorldEval, an action-centric discrete-diffusion world model that maps visual observations, language instructions, and action chunks into a shared unified token space and denoises them with a single self-attention backbone where actions function as first-class tokens. To realize reliable policy-world interaction, dWorldEval introduces a sparse keyframe memory that anchors global scene state while preserving fine-grained multi-view interaction cues, and leverages Progress-as-text to jointly generate future observations and success indicators. Extensive experiments on LIBERO, RoboTwin, and real-robot tasks demonstrate that dWorldEval significantly outperforms video diffusion baselines in action controllability, stabilizes long-horizon multi-view rollouts, enabling accurate policy ranking via automatic success estimation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d046895-05e5-4136-867d-e335da3bc943Builds on12
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang et al.NeurIPS 2025 · 949 citations
- Simple and Effective Masked Diffusion Language ModelsSubham S. Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan et al.NeurIPS 2024 · 929 citations
- Language Model Beats Diffusion - Tokenizer is key to visual generationLijun Yu, José Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari et al.ICLR 2024 · 609 citations
- Discrete Diffusion Modeling by Estimating the Ratios of the Data DistributionAaron Lou, Chenlin Meng, Stefano ErmonICML 2024 · 473 citations
Related papers
- Ctrl-World: A Controllable Generative World Model for Robot ManipulationYanjiang Guo, Lucy Xiaoyang Shi, Jianyu Chen, Chelsea FinnICLR 2026 · 163 citations
- Predicting What Matters: Robust Generalist Robot Policy Learning via Future Semantic MaskYunfan Lou, Xiaowei Chi, Xiaojie Zhang, Zezhong Qian et al.ICML 2026
- WorldGym: World Model as An Environment for Policy EvaluationJulian Quevedo, Ansh Kumar Sharma, Yixiang Sun, Varad Suryavanshi et al.ICLR 2026 · 58 citations
- DiffuView: Multi-View Diffusion Pretraining for 3D Aware Robotic ManipulationKaizhao Zhang, Tian Niu, Tianyu Liu, Chenen Guo et al.CVPR 2026
- IRASim: A Fine-Grained World Model for Robot ManipulationFangqi Zhu, Hongtao Wu, Song Guo, Yuxiao Liu et al.ICCV 2025 · 4 citations
