DiffCrossGait: Trajectory-Level Alignment for 2D-3D Cross-Modal Gait Recognition via Latent Diffusion
Zhiyang Lu, Ming Cheng
Abstract
Cross-modal 2D-3D gait recognition is impeded by inherent domain discrepancies between 2D silhouette and 3D LiDAR range-view representations. While prior methods align only final embeddings, we propose DiffCrossGait, which reformulates cross-modal matching as trajectorylevel alignment in an identity-relevant latent diffusion space, rather than assuming full equivalence between 2D and 3D observations. By driving both modalities with shared Gaussian noise within a latent space, we enable continuous alignment throughout the generative evolution. We introduce a Tri-Phase Alignment Strategy that exploits varying noise intensities to enforce identity anchoring, dynamics consistency, and cross-modal structural recoverability, thereby constraining both modalities to share denoising dynamics and bottleneck structure, which promotes modality-invariant gait features. Crucially, our framework decouples generative alignment from the discriminative backbone; the diffusion mechanism serves exclusively as a training objective, ensuring high inference efficiency by eliminating the computational overhead of iterative denoising. Extensive experiments on the SUSTech1K and FreeGait benchmarks demonstrate that DiffCross-Gait achieves state-of-the-art performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 201352b6-df3a-4306-811d-cdce93962a12Builds on25
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- RGB-Infrared Cross-Modality Person Re-Identification via Joint Pixel and Feature AlignmentGuan'an Wang, Tianzhu Zhang, Jian Cheng, Si Liu et al.ICCV 2019 · 464 citations
- Channel Augmented Joint Learning for Visible-Infrared RecognitionMang Ye, Weijian Ruan, Bo Du, Mike Zheng ShouICCV 2021 · 310 citations
Related papers
- Text-guided Feature Disentanglement for Cross-modal Gait RecognitionZhiyang Lu, Ming ChengCVPR 2026 · 1 citation
- X-Drive: Cross-modality Consistent Multi-Sensor Data Synthesis for Driving ScenariosYichen Xie, Chenfeng Xu, Chensheng Peng, Shuqi Zhao et al.ICLR 2025
- MS^2Gait: A Multi-Scale Spatio-Temporal Fusion Network for LiDAR-based Gait RecognitionShenyin Xu, Yishan Wang, Xinyu Li, Rui Liu et al.CVPR 2026
- Michelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent RepresentationZibo Zhao, Wen Liu, Xin Chen, Xianfang Zeng et al.NeurIPS 2023 · 279 citations
- SG-LDM: Semantic-Guided LiDAR Generation via Latent-Aligned DiffusionZhengkang Xiang, Zizhao Li, Amir Khodabandeh, Kourosh KhoshelhamICCV 2025
