Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
Moo Jin Kim, Yihuai Gao, Tsung-Yi Lin, Yen-Chen Lin, Yunhao Ge, Grace Lam, Percy Liang, Shuran Song, Ming-Yu Liu, Chelsea Finn, Jinwei Gu
摘要
Recent video generation models demonstrate remarkable ability to capture complex physical interactions and scene evolution over time. To leverage their spatiotemporal priors, robotics works have adapted video models for policy learning but introduce complexity by requiring multiple stages of post-training and new architectural components for action generation. In this work, we introduce Cosmos Policy, a simple approach for adapting a large pretrained video model (Cosmos-Predict2) into an effective robot policy through a single stage of post-training on the robot demonstration data collected on the target platform, with no architectural modifications. Cosmos Policy learns to directly generate robot actions encoded as latent frames within the video model's latent diffusion process, harnessing the model's pretrained priors and core learning algorithm to capture complex action distributions. Additionally, Cosmos Policy generates future state images and values (expected cumulative rewards), which are similarly encoded as latent frames, enabling test-time planning of action trajectories with higher likelihood of success. In our evaluations, Cosmos Policy achieves state-of-the-art performance on the LIBERO and RoboCasa simulation benchmarks (98.5% and 67.1% average success rates, respectively) and the highest average score in challenging real-world bimanual manipulation tasks, outperforming strong diffusion policies trained from scratch, video model-based policies, and state-of-the-art vision-language-action models fine-tuned on the same robot demonstrations. Furthermore, given policy rollout data, Cosmos Policy can learn from experience to refine its world model and value function and leverage model-based planning to achieve even higher success rates in challenging tasks. We release code, models, and training data at https://research.nvidia.com/labs/dir/cosmos-policy/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot LearningJiange Yang, Yansong Shi, Haoyi Zhu, Mingyu Liu 等CVPR 2026 · 被引用 47 次
- Evaluating Newtonian Mechanics in Video Generative Models with Real Physical SystemsAntonios Tragoudaras, Chenyu Zhang, Daniil Cherniavskii, Antonis Vozikis 等ICML 2026 · 被引用 39 次
- PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic ManipulationYuanzhe Liu, Jingyuan Zhu, Yuchen Mo, Gen Li 等CVPR 2026 · 被引用 31 次
- Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action ModelJohn Won, Kyungmin Lee, Huiwon Jang, Dongyoung Kim 等ICML 2026 · 被引用 22 次
- VideoGPA: Distilling Geometry Priors for 3D-Consistent Video GenerationHongyang Du, Hongyang Du, Xiaoyan Cong, Runhao Li 等ICML 2026 · 被引用 16 次
它引用的顶会 Paper12
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 被引用 1,170 次
- Temporal Difference Learning for Model Predictive ControlNicklas Hansen, Hao Su, Xiaolong WangICML 2022 · 被引用 388 次
相关 Paper
- ViPRA: Video Prediction for Robot ActionsSandeep Kumar Routray, Hengkai Pan, Unnat Jain, Shikhar Bahl 等ICLR 2026 · 被引用 30 次
- VideoVLA: Video Generators Can Be Generalizable Robot ManipulatorsYichao Shen, Fangyun Wei, Zhiying Du, Yaobo Liang 等NeurIPS 2025 · 被引用 73 次
- Video Prediction Policy: A Generalist Robot Policy with Predictive Visual RepresentationsYucheng Hu, Yanjiang Guo, Pengchao Wang, Xiaoyu Chen 等ICML 2025
- Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from VideosYi Chen, Yuying Ge, Weiliang Tang, Yizhuo Li 等ICCV 2025 · 被引用 5 次
- Learning to Act from Actionless Videos through Dense CorrespondencesPo-Chen Ko, Jiayuan Mao, Yilun Du, Shao-Hua Sun 等ICLR 2024 · 被引用 181 次
