World-In-World: World Models in a Closed-Loop World
Jiahan Zhang, Muqing Jiang, Nanru Dai, Taiming Lu, Arda Uzunoglu, Shunchi Zhang, Yana Wei, Jiahao Wang, Vishal M. Patel, Paul Pu Liang, Daniel Khashabi, Cheng Peng
摘要
Generative world models (WMs) can now simulate worlds with striking visual realism, which naturally raises the question of whether they can endow embodied agents with predictive perception for decision making. Progress on this question has been limited by fragmented evaluation: most existing benchmarks adopt open-loop protocols that emphasize visual quality in isolation, leaving the core issue of embodied utility unresolved, i.e., do WMs actually help agents succeed at embodied tasks? To address this gap, we introduce World-In-World, the first open platform that benchmarks WMs in a closed-loop setting that mirrors real agent-environment interactions. World-In-World provides a unified online planning strategy and a standardized action API, enabling heterogeneous WMs for decision making. We curate four closed-loop environments that rigorously evaluate diverse WMs, prioritize task success as the primary metric, and move beyond the common focus on visual quality; we also present the first data scaling law for world models in embodied settings. Our study uncovers three surprises: (1) visual quality alone does not guarantee task success—controllability matters more; (2) scaling post-training with action-observation data is more effective than upgrading the pretrained video generators; and (3) allocating more inference-time compute allows WMs to substantially improve closed-loop performance. By centering evaluation on closed-loop outcomes, World-In-World establishes a new benchmark for the systematic assessment of WMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Learning Latent Action World Models in the WildQuentin Garrido, Tushar Nagarajan, Basile Terver, Nicolas Ballas 等ICML 2026 · 被引用 38 次
- Goal Force: Teaching Video Models To Accomplish Physics-Conditioned GoalsNate Gillman, Yinghua Zhou, Zitian Tang, Evan Luo 等CVPR 2026 · 被引用 12 次
- Mode Seeking meets Mean Seeking for Fast Long Video GenerationShengqu Cai, Weili Nie, Chao Liu, Julius Berner 等ICML 2026 · 被引用 9 次
- Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous DrivingJiahao Wang, Bo Sun, Yijing Bai, Vincent Casser 等CVPR 2026 · 被引用 2 次
- DreamSAC: Learning Hamiltonian World Models via Symmetry ExplorationJinzhou Tang, Fan Feng, Minghao Fu, Wenjun Lin 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper48
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- MUSIQ: Multi-scale Image Quality TransformerJunjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar 等ICCV 2021 · 被引用 1,325 次
相关 Paper
- DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous DrivingYang Zhou, Hao Shao, Letian Wang, Zhuofan Zong 等ICLR 2026 · 被引用 20 次
- WorldLens: Full-Spectrum Evaluations of Driving World Models in Real WorldAo Liang, Lingdong Kong, Tianyi Yan, Hongsi Liu 等CVPR 2026 · 被引用 28 次
- iWorld-Bench: A Benchmark for Interactive World Models with a Unified Action Generation FrameworkJianjie Fang, Yingshan Lei, Qin Wan, Ziyou Wang 等ICML 2026 · 被引用 9 次
- Generative Visual Code Mobile World ModelsWoosung (Reiss) Koh, Sungjun Han, Segyu Lee, Se-Young Yun 等ICML 2026 · 被引用 6 次
- RLVR-World: Training World Models with Reinforcement LearningJialong Wu, Shaofeng Yin, Ningya Feng, Mingsheng LongNeurIPS 2025 · 被引用 52 次
