Human Motion Instruction Tuning
Lei Li, Sen Jia, Jianhao Wang, Zhongyu Jiang, Feng Zhou, Ju Dai, Tianfang Zhang, Zongkai Wu, Jenq-Neng Hwang
Abstract
This paper presents LLaMo (Large Language and Human Motion Assistant), a multimodal framework for human motion instruction tuning. In contrast to conventional instruction-tuning approaches that convert non-linguistic inputs, such as video or motion sequences, into language tokens, LLaMo retains motion in its native form for instruction tuning. This method preserves motion-specific details that are often diminished in tokenization, thereby improving the model's ability to interpret complex human behaviors. By processing both video and motion data alongside textual inputs, LLaMo enables a flexible, human-centric analysis. Experimental evaluations across high-complexity domains, including human behaviors and professional activities, indicate that LLaMo effectively captures domainspecific knowledge, enhancing comprehension and prediction in motion-intensive scenarios. We hope LLaMo offers a foundation for future multimodal AI systems with broad applications, from sports analytics to behavioral prediction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 055bb679-1149-4f53-8850-6b493df9ad8aCited by top-tier papers17
- Vision-Language-Action Pretraining from Large-Scale Human VideosHao Luo, Yicheng Feng, Wanpeng Zhang, Sipeng Zheng et al.ICML 2026 · 104 citations
- GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal GenerationTianchen Deng, Xuefeng Chen, Yi Chen, Qu Chen et al.CVPR 2026 · 31 citations
- ReTrack: Evidence-Driven Dual-Stream Directional Anchor Calibration Network for Composed Video RetrievalZixu Li, Yupeng Hu, Zhiwei Chen, Qinlei Huang et al.AAAI 2026 · 24 citations
- Intrinsic Entropy of Context Length Scaling in LLMsJingzhe Shi, Qinwei Ma, Hongyi Liu, Hang Zhao et al.ICLR 2026 · 17 citations
- Self-Destructive Language ModelsYuhui Wang, Rongyi Zhu, Ting WangICLR 2026 · 14 citations
Builds on13
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu et al.NeurIPS 2023 · 698 citations
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang et al.CVPR 2022 · 462 citations
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 279 citations
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui et al.EMNLP 2024 · 231 citations
Related papers
- Multiple Human Motion UnderstandingLei Li, Sen Jia, Jenq-Neng HwangAAAI 2026 · 4 citations
- LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive TokensZekun Li, Sizhe An, Chengcheng Tang, Chuan Guo et al.CVPR 2026 · 12 citations
- LLaMo: Large Language Model-based Molecular Graph AssistantJinyoung Park, Minseong Bae, Dohwan Ko, Hyunwoo J. KimNeurIPS 2024 · 33 citations
- LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of LivingDominick Reilly, Rajatsubhra Chakraborty, Arkaprava Sinha, Manish Kumar Govind et al.CVPR 2025
- CountLLM: Towards Generalizable Repetitive Action Counting via Large Language ModelZiyu Yao, Xuxin Cheng, Zhiqi Huang, Lei LiCVPR 2025
