GeoWorld: Geometric World Models
Zeyu Zhang, Danning Li, Ian Reid, Richard Hartley
Abstract
Energy-based predictive world models provide a powerful approach for multi-step visual planning by reasoning over latent energy landscapes rather than generating pixels. However, existing approaches face two major challenges: (i) their latent representations are typically learned in Euclidean space, neglecting the underlying geometric and hierarchical structure among states, and (ii) they struggle with long-horizon prediction, which leads to rapid degradation across extended rollouts. To address these challenges, we introduce GeoWorld, a geometric world model that preserves geometric structure and hierarchical relations through a Hyperbolic JEPA, which maps latent representations from Euclidean space onto hyperbolic manifolds. We further introduce Geometric Reinforcement Learning for energy-based optimization, enabling stable multi-step planning in hyperbolic latent space. Extensive experiments on CrossTask and COIN demonstrate around 3% SR improvement in 3-step planning and 2% SR improvement in 4-step planning compared to the state-of-the-art V-JEPA 2. Project website: https://steve-zeyu-zhang.github.io/GeoWorld.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6c3bc98f-ada2-4571-9c04-5dc887e11ebbBuilds on27
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised LearningAdrien Bardes, Jean Ponce, Yann LeCunICLR 2022 · 1,226 citations
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video DiffusionXun Huang, Zhengqi Li, Guande He, Mingyuan Zhou et al.NeurIPS 2025 · 628 citations
- LongLive: Real-time Interactive Long Video GenerationShuai Yang, Wei Huang, Ruihang Chu, Yicheng Xiao et al.ICLR 2026 · 241 citations
Related papers
- Temporal Straightening for Latent PlanningYing Wang, Oumayma Bounou, Gaoyue Zhou, Randall Balestriero et al.ICML 2026 · 19 citations
- GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA ManipulationJingjing Qian, Boyao Han, Chen Shi, Lei Xiao et al.CVPR 2026 · 19 citations
- Hyperbolic Deep Reinforcement LearningEdoardo Cetin, Benjamin Paul Chamberlain, Michael M. Bronstein, Jonathan J. HuntICLR 2023 · 4 citations
- Compositional Planning with Jumpy World ModelsJesse Farebrother, Matteo Pirotta, Andrea Tirinzoni, Marc Bellemare et al.ICML 2026 · 1 citation
- Long-Horizon Visual Planning with Goal-Conditioned Hierarchical PredictorsKarl Pertsch, Oleh Rybkin, Frederik Ebert, Shenghao Zhou et al.NeurIPS 2020 · 96 citations
