Bridging Gait Recognition and Large Language Models Sequence Modeling
Shaopeng Yang, Jilong Wang, Saihui Hou, Xu Liu, Chunshui Cao, Liang Wang, Yongzhen Huang
Abstract
Gait sequences exhibit sequential structures and contextual relationships similar to those in natural language, where each element-whether a word or a gait step-is connected to its predecessors and successors. This similarity enables the transformation of gait sequences into "texts" containing identity-related information. Large Language Models (LLMs), designed to understand and generate sequential data, can thus be utilized for gait sequence modeling to enhance gait recognition performance. Leveraging these insights, we make a pioneering effort to apply LLMs to gait recognition, which we refer to as GaitLLM. Specifically, we propose the Gait-to-Language (G2L) module, which converts gait sequences into a textual format suitable for LLMs, and the Language-to-Gait (L2G) module, which maps the LLM's output back to the gait feature space, thereby bridging the gap between LLM outputs and gait recognition. Notably, GaitLLM leverages the powerful modeling capabilities of LLMs without relying on complex architectural designs, improving gait recognition performance with only a small number of trainable parameters. Our method achieves state-of-the-art results on four popular gait datasets-SUSTech1K, CCPG, Gait3D, and GREW-demonstrating the effectiveness of applying LLMs in this domain. This work highlights the potential of LLMs to significantly enhance gait recognition, paving the way for future research and practical applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- BiggerGait: Unlocking Gait Recognition with Layer-wise Representations from Large Vision ModelsDingqiang Ye, Chao Fan, Zhanbo Huang, Chengwen Luo et al.NeurIPS 2025 · 28 citations
- EventGait: Towards Robust Gait Recognition with Event StreamsSenyan Xu, Shuai Chen, Chuanfu Shen, Kean Liu et al.CVPR 2026 · 2 citations
- MMGait: Towards Multi-Modal Gait RecognitionChenye Wang, Qingyuan Cai, Saihui Hou, Aoqi Li et al.CVPR 2026 · 1 citation
- Text-guided Feature Disentanglement for Cross-modal Gait RecognitionZhiyang Lu, Ming ChengCVPR 2026 · 1 citation
- DiffCrossGait: Trajectory-Level Alignment for 2D-3D Cross-Modal Gait Recognition via Latent DiffusionZhiyang Lu, Ming ChengICML 2026
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Gait Recognition via Effective Global-Local Feature Representation and Local Temporal AggregationBeibei Lin, Shunli Zhang, Xin YuICCV 2021 · 325 citations
Related papers
- Language-Guided and Motion-Aware Gait Representation for Generalizable RecognitionZhengxian Wu, Chuanrui Zhang, Shenao Jiang, Hangrui Xu et al.AAAI 2026 · 1 citation
- Vocabulary-Guided Gait RecognitionPanjian Huang, Saihui Hou, Chunshui Cao, Xu Liu et al.NeurIPS 2025 · 8 citations
- Translating Signals to Languages for sEMG-Based Activity RecognitionMing Wang, Haoxuan Qu, Qiuhong Ke, Wei Zhou et al.CVPR 2026
- SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity RecognitionZechen Li, Shohreh Deldari, Linyao Chen, Hao Xue et al.EMNLP 2025 · 9 citations
- BigGait: Learning Gait Representation You Want by Large Vision ModelsDingqiang Ye, Chao Fan, Jingzhe Ma, Xiaoming Liu et al.CVPR 2024 · 40 citations
