World-R1: Reinforcing 3D Constraints for Text-to-Video Generation
Weijie Wang, Xiaoxuan He, Youping Gu, Yifan Yang, Zeyu Zhang, Yefei He, Yanbo Ding, Xirui Hu, Donny Y. Chen, Zhiyuan He, Yuqing Yang, Bohan Zhuang
Abstract
2 Microsoft Research 3 Independent Researcher Recent video foundation models demonstrate impressive visual synthesis but frequently suffer from geometric inconsistencies. While existing methods attempt to inject 3D priors via architectural modifications, they often incur high computational costs and limit scalability. We propose World-R1, a framework that aligns video generation with 3D constraints through reinforcement learning. To facilitate this alignment, we introduce a specialized pure text dataset tailored for world simulation. Utilizing Flow-GRPO, we optimize the model using feedback from pre-trained 3D foundation models and vision-language models to enforce structural coherence without altering the underlying architecture. We further employ a periodic decoupled training strategy to balance rigid geometric consistency with dynamic scene fluidity. Extensive evaluations reveal that our approach significantly enhances 3D consistency while preserving the original visual quality of the foundation model, effectively bridging the gap between video generation and scalable world simulation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5c39c8a3-3996-491f-aa05-3690333da8deBuilds on32
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
- Depth Anything 3: Recovering the Visual Space from Any ViewsHaotong Lin, Sili Chen, Jun Hao Liew, Donny Y. Chen et al.ICLR 2026 · 720 citations
- Flow-GRPO: Training Flow Matching Models via Online RLJie Liu, Gongye Liu, Jiajun Liang, Yangguang Li et al.NeurIPS 2025 · 647 citations
Related papers
- WorldReel: 4D Video Generation with Consistent Geometry and Motion ModelingShaoheng Fang, Hanwen Jiang, Yunpeng Bai, Niloy J. Mitra et al.CVPR 2026 · 3 citations
- Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World ModelingHaoyu Wu, Diankun Wu, Tianyu He, Junliang Guo et al.ICLR 2026 · 89 citations
- WorldGrow: Generating Infinite 3D WorldSikuang Li, Chen Yang, Jiemin Fang, Taoran Yi et al.AAAI 2026 · 10 citations
- VideoGPA: Distilling Geometry Priors for 3D-Consistent Video GenerationHongyang Du, Hongyang Du, Xiaoyan Cong, Runhao Li et al.ICML 2026 · 16 citations
- World-consistent Video Diffusion with Explicit 3D ModelingQihang Zhang, Shuangfei Zhai, Miguel Ángel Bautista Martin, Kevin Miao et al.CVPR 2025
