Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs
Junhao Shi, Siyin Wang, Xiaopeng Yu, Li Ji, Jingjing Gong, Xipeng Qiu
摘要
Vision-Language-Action (VLA) models are fundamentally bottlenecked by the scarcity of expert demonstrations-triplets of observations, instructions, and actions that are costly to collect at scale. We argue that this bottleneck stems from conflating two distinct learning objectives: acquiring physical competence (how to move) and acquiring semantic alignment (what to do). Crucially, only the latter requires language supervision. Building on this Decomposition Hypothesis, we propose Task-Agnostic Pretraining (TAP), a two-stage framework that first learns transferable motor priors from cheap, unlabeled interaction data-including discarded off-task trajectories and autonomous robot play-via a self-supervised Inverse Dynamics objective. A lightweight second stage then grounds these priors in language using minimal expert data. On the SIMPLER benchmark, TAP matches models trained on over 1M expert trajectories while using orders of magnitude less labeled data, yielding a 10% absolute gain over standard behavior cloning. On a real-world WidowX platform, TAP retains 25% success under camera perturbations where internet-scale baselines collapse to 0%, demonstrating that taskagnostic pretraining produces robust, transferable physical representations and offers a scalable path forward for Embodied AI.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga 等NeurIPS 2022 · 被引用 458 次
- Unleashing Large-Scale Video Generative Pre-training for Visual Robot ManipulationHongtao Wu, Ya Jing, Chilam Cheang, Guangzeng Chen 等ICLR 2024 · 被引用 309 次
- Masked Autoencoding for Scalable and Generalizable Decision MakingFangchen Liu, Hao Liu, Aditya Grover, Pieter AbbeelNeurIPS 2022 · 被引用 63 次
- Inverse Dynamics Pretraining Learns Good Representations for Multitask ImitationDavid Brandfonbrener, Ofir Nachum, Joan BrunaNeurIPS 2023 · 被引用 38 次
- SMART: Self-supervised Multi-task pretrAining with contRol TransformersYanchao Sun, Shuang Ma, Ratnesh Madaan, Rogerio Bonatti 等ICLR 2023 · 被引用 4 次
相关 Paper
- Spatially Guided Training for Vision-Language-Action ModelJinhui Ye, Fangjing Wang, Ning Gao, Junqiu Yu 等ICLR 2026 · 被引用 6 次
- From Observation to Action: Latent Action-based Primitive Segmentation for VLA Pre-training in Industrial SettingsJiajie Zhang, Sören Schwertfeger, Alexander KleinerCVPR 2026
- TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic PoliciesRuijie Zheng, Yongyuan Liang, Shuaiyi Huang, Jianfeng Gao 等ICLR 2025
- Latent Action Pretraining from VideosSeonghyeon Ye, Joel Jang, Byeongguk Jeon, Se June Joo 等ICLR 2025
- LARA: Latent Action Representation Alignment for Vision-Language-Action ModelsMengya Liu, Baoxiong Jia, Jiangyong Huang, Jingze Zhang 等ICML 2026 · 被引用 3 次
