Inference-time Physics Alignment of Video Generative Models with Latent World Models
Jianhao Yuan, Xiaofeng Zhang, Felix Friedrich, Nicolas Beltran-Velez, Melissa Hall, Reyhane Askari Hemmat, Xiaochuang Han, Nicolas Ballas, Michal Drozdzal, Adriana Romero-Soriano
Abstract
State-of-the-art video generative models produce promising visual content yet often violate basic physics principles, limiting their utility. While some attribute this deficiency to insufficient physics understanding from pre-training, we find that the shortfall in physics plausibility also stems from suboptimal inference strategies. We therefore introduce WMReward and treat improving physics plausibility of video generation as an inference-time alignment problem. In particular, we leverage the strong physics prior of a latent world model (here, VJEPA-2) as a reward to search and steer multiple candidate denoising trajectories, enabling scaling test-time compute for better generation performance. Empirically, our approach substantially improves physics plausibility across image-conditioned, multiframe-conditioned, and text-conditioned generation settings, with validation from human preference study. Notably, in the ICCV 2025 Perception Test PhysicsIQ Challenge, we achieve a final score of 62.64%, winning first place and outperforming the previous state of the art by 7.42%. Our work demonstrates the viability of using latent world models to improve physics plausibility of video generation, beyond this specific instantiation or parameterization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 192b70c0-7a28-4fd5-93b5-a6a7a98479caCited by top-tier papers2
- Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases ThemWoojung Han, Seil Kang, Youngjun Jun, Min-Hung Chen et al.ICML 2026 · 2 citations
- MotiMotion: Motion-Controlled Video Generation with Visual ReasoningHsin-Ying Lee, Hanwen Jiang, Yiqun Mei, Jing Shi et al.ICML 2026
Builds on36
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- VideoPoet: A Large Language Model for Zero-Shot Video GenerationDan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama et al.ICML 2024 · 464 citations
- Universal Guidance for Diffusion ModelsArpit Bansal, Hong-Min Chu, Avi Schwarzschild, Soumyadip Sengupta et al.ICLR 2024 · 436 citations
- Learning Interactive Real-World SimulatorsSherry Yang, Yilun Du, Seyed Kamyar Seyed Ghasemipour, Jonathan Tompson et al.ICLR 2024 · 399 citations
Related papers
- PHANTOM: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical DynamicsYing Shen, Jerry Xiong, Tianjiao Yu, Ismini LourentzouCVPR 2026 · 12 citations
- PhysVid: Physics Aware Local Conditioning for Generative Video ModelsSaurabh Pathak, Elahe Arani, Mykola Pechenizkiy, Bahram ZonoozCVPR 2026 · 6 citations
- PhyCo: Learning Controllable Physical Priors for Generative MotionSriram Narayanan, Ziyu Jiang, Srinivasa G. Narasimhan, Manmohan ChandrakerCVPR 2026 · 6 citations
- PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff DropChenyu Li, Oscar Michel, Xichen Pan, Sainan Liu et al.ICML 2025
- Evaluating Newtonian Mechanics in Video Generative Models with Real Physical SystemsAntonios Tragoudaras, Chenyu Zhang, Daniil Cherniavskii, Antonis Vozikis et al.ICML 2026 · 39 citations
