CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning
Jiange Yang, Yansong Shi, Haoyi Zhu, Mingyu Liu, Kaijing Ma, Yating Wang, Gangshan Wu, Tong He, Limin Wang
Abstract
Unsupervised learning of latent motion from Internet videos is crucial for robot learning. Existing discrete methods generally mitigate the shortcut learning caused by extracting excessive static backgrounds through vector quantization with a small codebook size. However, they suffer from information loss and struggle to capture more complex and fine-grained dynamics. Moreover, there is an inherent gap between the distribution of discrete latent motion and continuous robot action, which hinders the joint learning of a unified policy. We propose CoMo, which aims to learn more precise continuous latent motion from internet-scale videos. CoMo employs an early temporal difference (Td) mechanism to increase the shortcut learning difficulty and explicitly enhance motion cues. Additionally, to ensure latent motion better captures meaningful foregrounds, we further propose a temporal contrastive learning (Tcl) scheme. Specifically, positive pairs are constructed with a small future frame temporal offset, while negative pairs are formed by directly reversing the temporal direction. The proposed Td and Tcl work synergistically and effectively ensure that the latent motion focuses better on the foreground and reinforces motion cues. Critically, CoMo exhibits strong zeroshot generalization, enabling it to generate effective pseudo action labels for unseen videos. Extensive simulated and real-world experiments show that policies co-trained with CoMo pseudo action labels achieve superior performance with both diffusion and auto-regressive architectures. The code will be available at https://github.com/MCG- NJU/CoMo.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 141f8d64-6f44-4d20-9551-b4e24cc8f96bCited by top-tier papers6
- Motus: A Unified Latent Action World ModelHongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang et al.CVPR 2026 · 271 citations
- StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State RepresentationMingyu Liu, Jiuhe Shu, Hui Chen, Zeju Li et al.CVPR 2026 · 14 citations
- Cross-Hand Latent Representation for Vision-Language-Action ModelsGuangqi Jiang, Yutong Liang, Jianglong Ye, Jia-Yang Huang et al.CVPR 2026 · 14 citations
- LAOF: Robust Latent Action Learning with Optical Flow ConstraintsXizhou Bu, Jiexi Lyu, Fulei Sun, Ruichen Yang et al.CVPR 2026 · 10 citations
- Bootstrap Dynamic-Aware 3D Visual Representation for Scalable Robot LearningQiwei Liang, Boyang Cai, Minghao Lai, Sitong Zhuang et al.CVPR 2026 · 6 citations
Builds on27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Autoregressive Image Generation without Vector QuantizationTianhong Li, Yonglong Tian, He Li, Mingyang Deng et al.NeurIPS 2024 · 758 citations
Related papers
- Unsupervised Pre-training for Temporal Action Localization TasksCan Zhang, Tianyu Yang, Junwu Weng, Meng Cao et al.CVPR 2022 · 56 citations
- Self-Supervised Video Representation Learning via Latent Time NavigationDi Yang, Yaohui Wang, Quan Kong, Antitza Dantcheva et al.AAAI 2023 · 18 citations
- From Pixels to Temporal Correlations: Learning Informative Representations for Reinforcement Learning Pre-trainingJinwen Wang, Youfang Lin, Xiaobo Hu, Siyu Yang et al.ACM MM 2025 · 1 citation
- Time-Equivariant Contrastive Video Representation LearningSimon Jenni, Hailin JinICCV 2021 · 64 citations
- Provable Ordering and Continuity in Vision-Language Pretraining for Generalizable Embodied AgentsZhizhen Zhang, Lei Zhu, Zhen Fang, Zi Huang et al.NeurIPS 2025 · 5 citations
