GSAlign: Geometric and Semantic Alignment Network for Aerial-Ground Person Re-Identification
Qiao Li, Jie Li, Yukang Zhang, Lei Tan, Jing Chen, Jiayi Ji
Abstract
Aerial-Ground person re-identification (AG-ReID) is an emerging yet challenging task that aims to match pedestrian images captured from drastically different viewpoints, typically from unmanned aerial vehicles (UAVs) and ground-based surveillance cameras. The task poses significant challenges due to extreme viewpoint discrepancies, occlusions, and domain gaps between aerial and ground imagery. While prior works have made progress by learning cross-view representations, they remain limited in handling severe pose variations and spatial misalignment. To address these issues, we propose a Geometric and Semantic Alignment Network (GSAlign) tailored for AG-ReID. GSAlign introduces two key components to jointly tackle geometric distortion and semantic misalignment in aerial-ground matching: a Learnable Thin Plate Spline (LTPS) Module and a Dynamic Alignment Module (DAM). The LTPS module adaptively warps pedestrian features based on a set of learned keypoints, effectively compensating for geometric variations caused by extreme viewpoint changes. In parallel, the DAM estimates visibility-aware representation masks that highlight visible body regions at the semantic level, thereby alleviating the negative impact of occlusions and partial observations in cross-view correspondence. A comprehensive evaluation on CARGO with four matching protocols demonstrates the effectiveness of GSAlign, achieving significant improvements of +18.8% in mAP and +16.8% in Rank-1 accuracy over previous state-of-the-art methods on the aerial-ground setting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e76f4b70-12c2-493c-a25b-efd1d3db450dBuilds on29
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- TransReID: Transformer-based Object Re-IdentificationShuting He, Hao Luo, Pichao Wang, Fan Wang et al.ICCV 2021 · 1,172 citations
- ABD-Net: Attentive but Diverse Person Re-IdentificationTianlong Chen, Shaojin Ding, Jingyi Xie, Ye Yuan et al.ICCV 2019 · 544 citations
- CLIP-ReID: Exploiting Vision-Language Model for Image Re-identification without Concrete Text LabelsSiyuan Li, Li Sun, Qingli LiAAAI 2023 · 355 citations
- Dynamic Prototype Mask for Occluded Person Re-IdentificationLei Tan, Pingyang Dai, Rongrong Ji, Yongjian WuACM MM 2022 · 94 citations
Related papers
- View-Aware Semantic Alignment for Aerial-Ground Person Re-IdentificationQuan Zhang, Zeqiang Cai, Peiming Zhao, Jingze Wu et al.CVPR 2026 · 1 citation
- View-decoupled Transformer for Person Re-identification under Aerial-ground Camera NetworkQuan Zhang, Lei Wang, Vishal M. Patel, Xiaohua Xie et al.CVPR 2024 · 30 citations
- AG-VPReID: A Challenging Large-Scale Benchmark for Aerial-Ground Video-based Person Re-IdentificationHuy Nguyen, Kien Nguyen, Akila Pemasiri, Feng Liu et al.CVPR 2025
- SeCap: Self-Calibrating and Adaptive Prompts for Cross-view Person Re-Identification in Aerial-Ground NetworksShining Wang, Yunlong Wang, Ruiqi Wu, Bingliang Jiao et al.CVPR 2025
- Bridging the Sky and Ground: Towards View-Invariant Feature Learning for Aerial-Ground Person Re-IdentificationWajahat Khalid, Bin Liu, Xulin Li, Muhammad Waqas et al.ICCV 2025 · 8 citations
