Vocabulary-Guided Gait Recognition
Panjian Huang, Saihui Hou, Chunshui Cao, Xu Liu, Yongzhen Huang
Abstract
What is a gait? Appearance-based gait networks consider a gait as the human shape and motion information from images. Model-based gait networks treat a gait as the human inherent structure from points. However, the considerations remain vague for humans to comprehend truly. In this work, we introduce a novel paradigm Vocabulary-Guided Gait Recognition, dubbed Gait-World, which attempts to explore gait concepts through human vocabularies with Vision-Language Models (VLMs). Although VLMs have achieved the remarkable progress in various vision tasks, the cognitive capability regarding gait modalities remains limited. The success element in Gait-World is the proper vocabulary prompt where this paradigm carefully selects gait cycle actions as Vocabulary Base, bridging the gait and vocabulary feature spaces and further promoting human understanding for the gait. How to extract gait features? Although previous gait networks have made significant progress, learning solely from gait modalities on limited gait databases makes it difficult to learn universal gait features for practicality. Therefore, we propose the first Gait-World model, dubbed -Gait, which guides the gait network learning with vocabulary knowledge from VLMs. However, due to the heterogeneity of the modalities, directly integrating vocabulary and gait features is highly challenging as they reside in different embedding spaces. To address the issues, -Gait designs Vocabulary Relation Mapper and Gait Fine grained Detector to map and establish vocabulary relations in the gait space for detecting corresponding gait features. Extensive experiments on CASIA-B, CCPG, SUSTech1K, Gait3D and GREW reveal the potential value and research directions of vocabulary information from VLMs in the gait field.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- EventGait: Towards Robust Gait Recognition with Event StreamsSenyan Xu, Shuai Chen, Chuanfu Shen, Kean Liu et al.CVPR 2026 · 2 citations
- MMGait: Towards Multi-Modal Gait RecognitionChenye Wang, Qingyuan Cai, Saihui Hou, Aoqi Li et al.CVPR 2026 · 1 citation
- BarbieGait: An Identity-Consistent Synthetic Human Dataset with Versatile Cloth-Changing for Gait RecognitionQingyuan Cai, Saihui Hou, Xuecai Hu, Yongzhen HuangCVPR 2026
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
- Gait Recognition via Effective Global-Local Feature Representation and Local Temporal AggregationBeibei Lin, Shunli Zhang, Xin YuICCV 2021 · 325 citations
Related papers
- Language-Guided and Motion-Aware Gait Representation for Generalizable RecognitionZhengxian Wu, Chuanrui Zhang, Shenao Jiang, Hangrui Xu et al.AAAI 2026 · 1 citation
- Bridging Gait Recognition and Large Language Models Sequence ModelingShaopeng Yang, Jilong Wang, Saihui Hou, Xu Liu et al.CVPR 2025
- BigGait: Learning Gait Representation You Want by Large Vision ModelsDingqiang Ye, Chao Fan, Jingzhe Ma, Xiaoming Liu et al.CVPR 2024 · 40 citations
- BiggerGait: Unlocking Gait Recognition with Layer-wise Representations from Large Vision ModelsDingqiang Ye, Chao Fan, Zhanbo Huang, Chengwen Luo et al.NeurIPS 2025 · 28 citations
- Splattalk: 3D VQA with Gaussian SplattingAnh Thai, Songyou Peng, Kyle Genova, Leonidas J. Guibas et al.ICCV 2025 · 4 citations
