Lune

CVPR2026Top-tier venue

Unlocking Motion from Large Vision Models with a Semantic and Kinematic Duality for Gait Recognition

Zhanbo Huang, Dingqiang Ye, Xiaoming Liu, Yu Kong

2026Year
4Citations
1Top-tier citations

Abstract

Existing set-based gait recognition methods achieve remarkable performance by capturing global semantic context. However, their order-invariant nature prevents them from modeling the fine-grained kinematic patterns that unfold over time. To unify the global and process-level representations, we propose GaitMax, a framework capturing both semantic context and kinematic motion. GaitMax leverages attention-based spatiotemporal modeling to dynamically represent detailed part-level trajectories. While this detailed representation is more powerful, it also captures more nuisance factors (e.g., clothing, viewpoint), leading to potential shortcuts. To mitigate this, we introduce Conditional Decorrelation Loss (CDLoss), which explicitly disentangles the gait embeddings from nuisance factors using vision-language supervision. This loss requires high-quality nuisance descriptions. We therefore construct GCaption, a new resource that provides natural language annotations for multiple gait datasets, moving beyond simple categorical labels. GCaption not only enables our CDLoss but also serves as a foundation for future context-aware gait analysis. Models, code, and resources are available at https://zbhuang.com/gait-max.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext b2b962d0-d50f-42b4-b4ec-83d050e874e4

Cited by top-tier papers1

Ask how each one uses it

Builds on20

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines