Gait Transformer: End-to-End Transformer Backbone for Gait Recognition
Saihui Hou, Wenpeng Lang, Jilong Wang, Yan Huang, Liang Wang, Yongzhen Huang
Abstract
Gait recognition has emerged as a promising biometric technique for long-distance and non-intrusive human identification. While Transformers have revolutionized vision tasks, their adaptation to gait recognition remains underexplored due to domain-specific challenges such as sparse silhouette modality, spatial-temporal dynamics, fine-grained motion cues, and limited training data. In this paper, we propose Gait Transformer (GaT), an end-to-end Transformer backbone specifically tailored for silhouette-based gait recognition. GaT introduces three key components: (1) a hybrid patch embedding module that combines convolutional stems with group-batch normalization to enhance structural preservation; (2) a decomposed token mixer that explicitly models both short-range and long-range dependencies across spatial-temporal dimensions; and (3) a hybrid positional encoding strategy that integrates absolute, relative, and rotary embeddings to support efficient training under data scarcity. Without relying on any pretraining, GaT achieves state-of-the-art performance on Gait3D, GREW, and CCGR-MINI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Segmenter: Transformer for Semantic SegmentationRobin Strudel, Ricardo Garcia, Ivan Laptev, Cordelia SchmidICCV 2021 · 1,898 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- Early Convolutions Help Transformers See BetterTete Xiao, Mannat Singh, Eric Mintun, Trevor Darrell et al.NeurIPS 2021 · 974 citations
Related papers
- Multi-modal Gait Recognition via Effective Spatial-Temporal Feature FusionYufeng Cui, Yimei KangCVPR 2023
- Learning Visual Prompt for Gait RecognitionKang Ma, Ying Fu, Chunshui Cao, Saihui Hou et al.CVPR 2024 · 24 citations
- HybridGait: A Benchmark for Spatial-Temporal Cloth-Changing Gait Recognition with Hybrid ExplorationsYilan Dong, Chunlin Yu, Ruiyang Ha, Ye Shi et al.AAAI 2024 · 31 citations
- GLGait: A Global-Local Temporal Receptive Field Network for Gait Recognition in the WildGuozhen Peng, Yunhong Wang, Yuwei Zhao, Shaoxiong Zhang et al.ACM MM 2024 · 12 citations
- On Model and Data Scaling for Skeleton-based Self-Supervised Gait RecognitionAdrian Cosma, Andy-Eduard Catruna, Emilian RadoiAAAI 2026 · 1 citation
