General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling
Huaihai Lyu, Chaofan Chen, Mingyu Cao, Yuheng Ji, Changsheng Xu
Abstract
Achieving robust generalization from limited data is a central challenge in embodied intelligence. Prevailing methods fail by regressing absolute coordinates, which violates the principle of general covariance. Fundamentally, this conflates the intrinsic task geometry with rigid execution patterns, binding policies to specific motion styles and fixed speeds. To resolve this, we propose the Generalized Action Manifold (GAM) framework that enforces general covariance through structural disentanglement. Specifically, GAM realizes the manifold by enforcing invariance across two orthogonal dimensions: (1) Temporal Invariance, utilizing an Arc-Length Parameterizer to orthogonalize the spatial path geometry from temporal dynamics, ensuring robustness to velocity variations; (2) Geometric Invariance, where a Schema-Affine-Factorization mechanism maps trajectories to canonical "world lines" in a pose-normalized coordinate frame. This distinguishes invariant geometric schemas from affine modulations, ensuring spatial generalizability. By integrating GAM within a structured Vision-Language-Action (VLA) architecture, we enable sparse demonstrations to densely populate a continuous, valid action manifold. Empirical results demonstrate that GAM enables superior transfer and robustness capabilities, outperforming geometry-agnostic baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Behavior Generation with Latent ActionsSeungjae Lee, Yibin Wang, Haritheja Etukuru, H. Jin Kim et al.ICML 2024 · 154 citations
- BEAST: Efficient Tokenization of B-Splines Encoded Action Sequences for Imitation LearningHongyi Zhou, Weiran Liao, Xi Huang, Yucheng Tang et al.NeurIPS 2025 · 32 citations
- Imitating Human Behaviour with Diffusion ModelsTim Pearce, Tabish Rashid, Anssi Kanervisto, David Bignell et al.ICLR 2023 · 23 citations
Related papers
- Spatially Guided Training for Vision-Language-Action ModelJinhui Ye, Fangjing Wang, Ning Gao, Junqiu Yu et al.ICLR 2026 · 6 citations
- AIR-VLA: Vision-Language-Action Systems for Aerial ManipulationJianli Sun, Bin Tian, Qiyao Zhang, Chengxiang Li et al.ICML 2026 · 4 citations
- EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric FlowYixiang Chen, Peiyan Li, Yan Huang, Jiabing Yang et al.ICCV 2025 · 2 citations
- Video2Robo: 3DGS-based Synthetic Data from One Video Enables Scalable Robot LearningYinan Deng, Kejia Hu, Ye Chen, Jianyu Dou et al.CVPR 2026
- LAST: Bridging Vision-Language and Action Manifolds via Gromov-Wasserstein AlignmentHuaihai Lyu, Chaofan Chen, Yuheng Ji, Xiansheng Chen et al.ICML 2026 · 1 citation
