Selective and Orthogonal Feature Activation for Pedestrian Attribute Recognition
Junyi Wu, Yan Huang, Min Gao, Yuzhen Niu, Mingjing Yang, Zhipeng Gao, Jianqiang Zhao
摘要
Pedestrian Attribute Recognition (PAR) involves identifying the attributes of individuals in person images. Existing PAR methods typically rely on CNNs as the backbone network to extract pedestrian features. However, CNNs process only one adjacent region at a time, leading to the loss of long-range inter-relations between different attribute-specific regions. To address this limitation, we leverage the Vision Transformer (ViT) instead of CNNs as the backbone for PAR, aiming to model long-range relations and extract more robust features. However, PAR suffers from an inherent attribute imbalance issue, causing ViT to naturally focus more on attributes that appear frequently in the training set and ignore some pedestrian attributes that appear less. The native features extracted by ViT are not able to tolerate the imbalance attribute distribution issue. To tackle this issue, we propose two novel components: the Selective Feature Activation Method (SFAM) and the Orthogonal Feature Activation Loss. SFAM smartly suppresses the more informative attribute-specific features, compelling the PAR model to capture discriminative features from regions that are easily overlooked. The proposed loss enforces an orthogonal constraint on the original feature extracted by ViT and the suppressed features from SFAM, promoting the complementarity of features in space. We conduct experiments on several benchmark PAR datasets, including PETA, PA100K, RAPv1, and RAPv2, demonstrating the effectiveness of our method. Specifically, our method outperforms existing state-of-the-art approaches by GRL, IAA-Caps, ALM, and SSC in terms of mA on the four datasets, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Pedestrian Attribute Recognition: A New Benchmark Dataset and a Large Language Model Augmented FrameworkJiandong Jin, Xiao Wang, Qian Zhu, Haiyang Wang 等AAAI 2025 · 被引用 19 次
- Joint Implicit and Explicit Language Learning for Pedestrian Attribute RecognitionYukang Zhang, Lei Tan, Yang Lu, Yan Yan 等AAAI 2026 · 被引用 1 次
- Enhanced Visual-Semantic Interaction with Tailored Prompts for Pedestrian Attribute RecognitionJunyi Wu, Yan Huang, Min Gao, Yuzhen Niu 等CVPR 2025
- Language-Guided Audio-Visual Learning for Long-Term Sports AssessmentHuangbiao Xu, Xiao Ke, Huanqi Wu, Rui Xu 等CVPR 2025
它引用的顶会 Paper13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- TransReID: Transformer-based Object Re-IdentificationShuting He, Hao Luo, Pichao Wang, Fan Wang 等ICCV 2021 · 被引用 1,172 次
- Towards a Unified Middle Modality Learning for Visible-Infrared Person Re-IdentificationYukang Zhang, Yan Yan, Yang Lu, Hanzi WangACM MM 2021 · 被引用 219 次
- Improving Pedestrian Attribute Recognition With Weakly-Supervised Multi-Scale Attribute-Specific LocalizationChufeng Tang, Lu Sheng, Zhaoxiang Zhang, Xiaolin HuICCV 2019 · 被引用 153 次
- Clothing Status Awareness for Long-Term Person Re-IdentificationYan Huang, Qiang Wu, Jingsong Xu, Yi Zhong 等ICCV 2021 · 被引用 132 次
相关 Paper
- POAR: Towards Open Vocabulary Pedestrian Attribute RecognitionYue Zhang, Suchen Wang, Shichao Kan, Zhenyu Weng 等ACM MM 2023 · 被引用 11 次
- Learning Disentangled Attribute Representations for Robust Pedestrian Attribute RecognitionJian Jia, Naiyu Gao, Fei He, Xiaotang Chen 等AAAI 2022 · 被引用 50 次
- Spatial and Semantic Consistency Regularizations for Pedestrian Attribute RecognitionJian Jia, Xiaotang Chen, Kaiqi HuangICCV 2021 · 被引用 80 次
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski 等AAAI 2022 · 被引用 529 次
- Pose-guided Inter- and Intra-part Relational Transformer for Occluded Person Re-IdentificationZhongxing Ma, Yifan Zhao, Jia LiACM MM 2021 · 被引用 66 次
