Text-guided Feature Disentanglement for Cross-modal Gait Recognition
Zhiyang Lu, Ming Cheng
Abstract
Gait recognition is a biometric technique that identifies individuals based on their walking patterns, offering advantages in long-range, non-intrusive scenarios. However, real-world scenarios often involve heterogeneous sensing modalities such as LiDAR and RGB cameras, making LiDAR-Camera Cross-modal Gait recognition (LCCGR) a critical yet challenging task due to the substantial modality gap between 2D videos and 3D point cloud sequences. To address this challenge, we propose TCFDNet, a Text-guided Cross-modal Feature Disentanglement Network, which leverages modality-aware textual priors as semantic anchors to guide the learning of disentangled modality-shared representations. Specifically, we construct a Gait Modality Text Dictionary (GMTD) using large language models to generate rich semantic descriptions of gait across modalities and viewpoints. A CLIP-based Multi-grained Feature Encoder then aligns visual and textual features within a unified vision-language space. Furthermore, the Text-guided Feature Disentanglement (TFD) module selects the topk matched textual descriptions to reconstruct modality-specific representations and derive modality-shared features via residual decomposition and orthogonality constraints. To mitigate the fragility of the disentangled shared features, we propose a Feature Stability Enhancement (FSE) module, which models spatial and channel-wise correlations to improve feature robustness. In addition, a cross-modal patch exchange strategy is introduced to further improve generalization. Extensive experiments on SUSTech1K and FreeGait datasets demonstrate that TCFDNet achieves new state-of-the-art results and validate the effectiveness of the proposed modules.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng et al.EMNLP 2024 · 479 citations
- Gait Recognition via Effective Global-Local Feature Representation and Local Temporal AggregationBeibei Lin, Shunli Zhang, Xin YuICCV 2021 · 325 citations
- Channel Augmented Joint Learning for Visible-Infrared RecognitionMang Ye, Weijian Ruan, Bo Du, Mike Zheng ShouICCV 2021 · 310 citations
Related papers
- Walking Further: Semantic-Aware Multimodal Gait Recognition Under Long-Range ConditionsZhiyang Lu, Wen Jiang, Tianren Wu, Zhichao Wang et al.AAAI 2026
- DiffCrossGait: Trajectory-Level Alignment for 2D-3D Cross-Modal Gait Recognition via Latent DiffusionZhiyang Lu, Ming ChengICML 2026
- LidarGait: Benchmarking 3D Gait Recognition with Point CloudsChuanfu Shen, Fan Chao, Wei Wu, Rui Wang et al.CVPR 2023
- Gait Recognition in Large-scale Free Environment via Single LiDARXiao Han, Yiming Ren, Peishan Cong, Yujing Sun et al.ACM MM 2024 · 10 citations
- GTR-Loc: Geospatial Text Regularization Assisted Outdoor LiDAR LocalizationShangshu Yu, Wen Li, Xiaotian Sun, Zhimin Yuan et al.NeurIPS 2025 · 2 citations
