World Guidance: World Modeling in Condition Space for Action Generation
Yue Su, Sijin Chen, Haixin Shi, Mingyu Liu, Zhengshen Zhang, Ningyuan Huang, Weiheng Zhong, Zhengbang Zhu, Yuxiao Liu, Xihui Liu
Abstract
Leveraging future observation modeling to facilitate action generation presents a promising avenue for enhancing the capabilities of Vision-Language-Action (VLA) models. However, existing approaches struggle to strike a balance between maintaining efficient, predictable future representations and preserving sufficient fine-grained information to guide precise action generation. To address this limitation, we propose WoG (World Guidance), a framework that maps future observations into compact conditions by injecting them into the action inference pipeline. The VLA is then trained to simultaneously predict these compressed conditions alongside future actions, thereby achieving effective world modeling within the condition space for action inference. We demonstrate that modeling and predicting this condition space not only facilitates fine-grained action generation but also exhibits superior generalization capabilities. Moreover, it learns effectively from substantial human manipulation videos. Extensive experiments across both simulation and real-world environments validate that WoG significantly outperforms existing methods based on future prediction. Project page is available at: https://selen-suyue.github.io/WoGNet/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8eac25e7-2be0-43d3-a365-8b1413f47191Cited by top-tier papers1
Ask how each one uses itBuilds on26
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language ModelsSiddharth Karamcheti, Suraj Nair, Ashwin Balakrishna, Percy Liang et al.ICML 2024 · 306 citations
- DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World KnowledgeWenyao Zhang, Hongsi Liu, Zekun Qi, Yunnan Wang et al.NeurIPS 2025 · 244 citations
Related papers
- WorldAgen: Unified State-Action Prediction with Test-Time World Model TrainingChi Wan, Kangrui Wang, Yuan Si, Pingyue Zhang et al.AAAI 2026
- Chain of World: World Model Thinking in Latent MotionFuxiang Yang, Donglin Di, Lulu Tang, Xuancheng Zhang et al.CVPR 2026 · 11 citations
- DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous DrivingYingyan Li, Shuyao Shang, Weisong Liu, Bing Zhan et al.ICLR 2026 · 134 citations
- Improving Vision-and-Language Navigation by Generating Future-View Image SemanticsJialu Li, Mohit BansalCVPR 2023
- Unified Vision-Language-Action ModelYuqi Wang, Xinghang Li, Wenxuan Wang, Junbo Zhang et al.ICLR 2026 · 144 citations
