Multi-modal Gait Recognition via Effective Spatial-Temporal Feature Fusion
Yufeng Cui, Yimei Kang
Abstract
Gait recognition is a biometric technology that identifies people by their walking patterns. The silhouettesbased method and the skeletons-based method are the two most popular approaches. However, the silhouette data are easily affected by clothing occlusion, and the skeleton data lack body shape information. To obtain a more robust and comprehensive gait representation for recognition, we propose a transformer-based gait recognition framework called MMGaitFormer, which effectively fuses and aggregates the spatial-temporal information from the skeletons and silhouettes. Specifically, a Spatial Fusion Module (SFM) and a Temporal Fusion Module (TFM) are proposed for effective spatial-level and temporal-level feature fusion, respectively. The SFM performs fine-grained body parts spatial fusion and guides the alignment of each part of the silhouette and each joint of the skeleton through the attention mechanism. The TFM performs temporal modeling through Cycle Position Embedding (CPE) and fuses temporal information of two modalities. Experiments demonstrate that our MMGaitFormer achieves state-of-the-art performance on popular gait datasets. For the most challenging "CL" (i.e., walking in different clothes) condition in CASIA-B, our method achieves a rank-1 accuracy of 94.8%, which outperforms the state-of-the-art single-modal methods by a large margin.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1fcd0213-6ead-47fa-bb05-d0bc5916154eCited by top-tier papers11
- It Takes Two: Accurate Gait Recognition in the Wild via Cross-granularity AlignmentJinkai Zheng, Xinchen Liu, Boyue Zhang, Chenggang Yan et al.ACM MM 2024 · 14 citations
- Vocabulary-Guided Gait RecognitionPanjian Huang, Saihui Hou, Chunshui Cao, Xu Liu et al.NeurIPS 2025 · 8 citations
- GaitSnippet: Gait Recognition Beyond Unordered Sets and Ordered SequencesSaihui Hou, Chenye Wang, Wenpeng Lang, Zhengxiang Lan et al.ICLR 2026 · 5 citations
- WaveLoss: An Adaptive Dynamic Loss for Deep Gait RecognitionZicheng Wang, Qiuxia WuAAAI 2025 · 4 citations
- DepthGait: Multi-Scale Cross-Level Feature Fusion of RGB-Derived Depth and Silhouette Sequences for Robust Gait RecognitionXinzhu Li, Juepeng Zheng, Yikun Chen, Xudong Mao et al.ACM MM 2025 · 2 citations
Builds on4
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Gait Recognition via Effective Global-Local Feature Representation and Local Temporal AggregationBeibei Lin, Shunli Zhang, Xin YuICCV 2021 · 325 citations
- Gait Recognition with Multiple-Temporal-Scale 3D Convolutional Neural NetworkBeibei Lin, Shunli Zhang, Feng BaoACM MM 2020 · 173 citations
- GaitPart: Temporal Part-Based Model for Gait RecognitionChao Fan, Yunjie Peng, Chunshui Cao, Xu Liu et al.CVPR 2020
Related papers
- GaitCycFormer: Leveraging Gait Cycles and Transformers for Gait Emotion RecognitionQingyang Zeng, Lin ShangAAAI 2025 · 6 citations
- HybridGait: A Benchmark for Spatial-Temporal Cloth-Changing Gait Recognition with Hybrid ExplorationsYilan Dong, Chunlin Yu, Ruiyang Ha, Ye Shi et al.AAAI 2024 · 31 citations
- Gait Transformer: End-to-End Transformer Backbone for Gait RecognitionSaihui Hou, Wenpeng Lang, Jilong Wang, Yan Huang et al.AAAI 2026
- DyGait: Exploiting Dynamic Representations for High-performance Gait RecognitionMing Wang, Xianda Guo, Beibei Lin, Tian Yang et al.ICCV 2023 · 81 citations
- Hierarchical Spatio-Temporal Representation Learning for Gait RecognitionLei Wang, Bo Liu, Fangfang Liang, Bincheng WangICCV 2023 · 43 citations
