Skeleton-in-Context: Unified Skeleton Sequence Modeling with In-Context Learning
Xinshun Wang, Zhongbin Fang, Xia Li, Xiangtai Li, Chen Chen, Mengyuan Liu
Abstract
In-context learning provides a new perspective for multi-task modeling for vision and NLP. Under this setting, the model can perceive tasks from prompts and accomplish them without any extra task-specific head predictions or model fine-tuning. However, skeleton sequence modeling via in-context learning remains unexplored. Directly applying existing in-context models from other areas onto skeleton sequences fails due to the similarity between inter-frame and cross-task poses, which makes it exceptionally hard to perceive the task correctly from a subtle context. To address this challenge, we propose Skeleton-in-Context (SiC), an effective framework for in-context skeleton sequence modeling. Our SiC is able to handle multiple skeleton-based tasks simultaneously after a single training process and accomplish each task from context according to the given prompt. It can further generalize to new, unseen tasks according to customized prompts. To facilitate context perception, we additionally propose a task-unified prompt, which adaptively learns tasks of different natures, such as partial joint-level generation, sequence-level prediction, or 2D-to-3D motion prediction. We conduct extensive experiments to evaluate the effectiveness of our SiC on multiple tasks, including motion prediction, pose estimation, joint completion, and future pose estimation. We also evaluate its generalization capability on unseen tasks such as motion-in-between. These experiments show that our model achieves state-of-the-art multi-task performance and even outperforms single-task methods on certain tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 682b4779-1e35-4c2b-81f1-2b5853a0edabCited by top-tier papers18
- OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and UnderstandingTao Zhang, Xiangtai Li, Hao Fei, Haobo Yuan et al.NeurIPS 2024 · 186 citations
- Point Cloud Mamba: Point Cloud Learning via State Space ModelTao Zhang, Haobo Yuan, Lu Qi, Jiangning Zhang et al.AAAI 2025 · 110 citations
- CSBrain: A Cross-scale Spatiotemporal Brain Foundation Model for EEG DecodingYuchen Zhou, Jiamin Wu, Zichen Ren, Zhouheng Yao et al.NeurIPS 2025 · 71 citations
- TCPFormer: Learning Temporal Correlation with Implicit Pose Proxy for 3D Human Pose EstimationJiajie Liu, Mengyuan Liu, Hong Liu, Wenhao LiAAAI 2025 · 27 citations
- USDRL: Unified Skeleton-Based Dense Representation Learning with Multi-Grained Feature DecorrelationWanjiang Weng, Hongsong Wang, Junbo Wang, Lei He et al.AAAI 2025 · 15 citations
Builds on36
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
Related papers
- SkeletonContext: Skeleton-side Context Prompt Learning for Zero-Shot Skeleton-based Action RecognitionNing Wang, Tieyue Wu, Naeha Sharif, Farid Boussaïd et al.CVPR 2026 · 3 citations
- Prompted Contrast with Masked Motion Modeling: Towards Versatile 3D Action Representation LearningJiahang Zhang, Lilang Lin, Jiaying LiuACM MM 2023 · 26 citations
- Superman: Unifying Skeleton and Vision for Human Motion Perception and GenerationXinshun Wang, Peiming Li, Ziyi Wang, Zhongbin Fang et al.CVPR 2026
- In-Context Learning Unlocked for Diffusion ModelsZhendong Wang, Yifan Jiang, Yadong Lu, Yelong Shen et al.NeurIPS 2023 · 128 citations
- MS2L: Multi-Task Self-Supervised Learning for Skeleton Based Action RecognitionLilang Lin, Sijie Song, Wenhan Yang, Jiaying LiuACM MM 2020 · 217 citations
