Lune

ICML2026Top-tier venue

Dyn-VPP: Video Prediction Policy Optimization for Improved Visual Dynamics

Zirui Ge, Pengxiang Ding, BaoHuaYin, Yemin Wang, Qishen Wang, Zhiyong Xie, Hengtao Li, Runze Suo, Wenxuan Song, Han Zhao, Shangke Lyu, Haoang Li

2026Year

Abstract

Video action models are a promising foundation for Vision-Language-Action (VLA) because they can learn rich visual dynamics directly from video. However, likelihood-oriented training of diffusion predictors emphasizes globally plausible futures and does not guarantee precision-critical visual dynamics needed for manipulation, so small prediction errors can be amplified by downstream policies. We propose Dyn-VPP, a posttraining framework that casts multi-step denoising as policy optimization and aligns predicted future latents with expert visual dynamics via a verifiable terminal reward, without modifying any architecture. This enables explicit optimization of dynamics signals that are not captured by likelihood-only training. As a result, Dyn-VPP yields more accurate visual dynamics and improves downstream task execution. Experiments across diverse simulated and real-world manipulation settings show improved dynamics consistency and consistently higher task success.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 68f47872-72da-43e5-9499-1b48e5860c7b

Builds on22

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines