Selective and Orthogonal Feature Activation for Pedestrian Attribute Recognition
Junyi Wu, Yan Huang, Min Gao, Yuzhen Niu, Mingjing Yang, Zhipeng Gao, Jianqiang Zhao
Abstract
Pedestrian Attribute Recognition (PAR) involves identifying the attributes of individuals in person images. Existing PAR methods typically rely on CNNs as the backbone network to extract pedestrian features. However, CNNs process only one adjacent region at a time, leading to the loss of long-range inter-relations between different attribute-specific regions. To address this limitation, we leverage the Vision Transformer (ViT) instead of CNNs as the backbone for PAR, aiming to model long-range relations and extract more robust features. However, PAR suffers from an inherent attribute imbalance issue, causing ViT to naturally focus more on attributes that appear frequently in the training set and ignore some pedestrian attributes that appear less. The native features extracted by ViT are not able to tolerate the imbalance attribute distribution issue. To tackle this issue, we propose two novel components: the Selective Feature Activation Method (SFAM) and the Orthogonal Feature Activation Loss. SFAM smartly suppresses the more informative attribute-specific features, compelling the PAR model to capture discriminative features from regions that are easily overlooked. The proposed loss enforces an orthogonal constraint on the original feature extracted by ViT and the suppressed features from SFAM, promoting the complementarity of features in space. We conduct experiments on several benchmark PAR datasets, including PETA, PA100K, RAPv1, and RAPv2, demonstrating the effectiveness of our method. Specifically, our method outperforms existing state-of-the-art approaches by GRL, IAA-Caps, ALM, and SSC in terms of mA on the four datasets, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Pedestrian Attribute Recognition: A New Benchmark Dataset and a Large Language Model Augmented FrameworkJiandong Jin, Xiao Wang, Qian Zhu, Haiyang Wang et al.AAAI 2025 · 19 citations
- Joint Implicit and Explicit Language Learning for Pedestrian Attribute RecognitionYukang Zhang, Lei Tan, Yang Lu, Yan Yan et al.AAAI 2026 · 1 citation
- Enhanced Visual-Semantic Interaction with Tailored Prompts for Pedestrian Attribute RecognitionJunyi Wu, Yan Huang, Min Gao, Yuzhen Niu et al.CVPR 2025
- Language-Guided Audio-Visual Learning for Long-Term Sports AssessmentHuangbiao Xu, Xiao Ke, Huanqi Wu, Rui Xu et al.CVPR 2025
Builds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- TransReID: Transformer-based Object Re-IdentificationShuting He, Hao Luo, Pichao Wang, Fan Wang et al.ICCV 2021 · 1,172 citations
- Towards a Unified Middle Modality Learning for Visible-Infrared Person Re-IdentificationYukang Zhang, Yan Yan, Yang Lu, Hanzi WangACM MM 2021 · 219 citations
- Improving Pedestrian Attribute Recognition With Weakly-Supervised Multi-Scale Attribute-Specific LocalizationChufeng Tang, Lu Sheng, Zhaoxiang Zhang, Xiaolin HuICCV 2019 · 153 citations
- Clothing Status Awareness for Long-Term Person Re-IdentificationYan Huang, Qiang Wu, Jingsong Xu, Yi Zhong et al.ICCV 2021 · 132 citations
Related papers
- POAR: Towards Open Vocabulary Pedestrian Attribute RecognitionYue Zhang, Suchen Wang, Shichao Kan, Zhenyu Weng et al.ACM MM 2023 · 11 citations
- Learning Disentangled Attribute Representations for Robust Pedestrian Attribute RecognitionJian Jia, Naiyu Gao, Fei He, Xiaotang Chen et al.AAAI 2022 · 50 citations
- Spatial and Semantic Consistency Regularizations for Pedestrian Attribute RecognitionJian Jia, Xiaotang Chen, Kaiqi HuangICCV 2021 · 80 citations
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski et al.AAAI 2022 · 529 citations
- Pose-guided Inter- and Intra-part Relational Transformer for Occluded Person Re-IdentificationZhongxing Ma, Yifan Zhao, Jia LiACM MM 2021 · 66 citations
