POAR: Towards Open Vocabulary Pedestrian Attribute Recognition
Yue Zhang, Suchen Wang, Shichao Kan, Zhenyu Weng, Yigang Cen, Yap-Peng Tan
Abstract
Pedestrian attribute recognition (PAR) aims to predict the attributes of a target pedestrian. Recent methods often address the PAR problem by training a multi-label classifier with predefined attribute classes, but they can hardly exhaust all possible pedestrian attributes in the real world. To tackle this problem, we propose a novel Pedestrian Open-Attribute Recognition (POAR) approach by formulating the problem as a task of image-text search. Our approach employs a Transformer-based Encoder with a Masking Strategy (TEMS) to focus on the attributes of specific pedestrian parts (e.g., head, upper body, lower body, feet, etc.), and introduces a set of attribute tokens to encode the corresponding attributes into visual embeddings. Each attribute category is described as a natural language sentence and encoded by the text encoder. Then, we compute the similarity between the visual and text embeddings to find the best attribute descriptions for the input images. To handle multiple attributes of a single pedestrian, we propose a Many-To-Many Contrastive (MTMC) loss with masked tokens. In addition, we propose a Grouped Knowledge Distillation (GKD) method to minimize the disparity between visual embeddings and unseen attribute text embeddings. We evaluate our proposed method on three PAR datasets with an open-attribute setting. The results demonstrate the effectiveness of our method as a strong baseline for the POAR task. Our code is available at https://github.com/IvyYZ/POAR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e5240409-37cb-4c74-9b82-a772080db476Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
- TransReID: Transformer-based Object Re-IdentificationShuting He, Hao Luo, Pichao Wang, Fan Wang et al.ICCV 2021 · 1,172 citations
- Open-Set Recognition: A Good Closed-Set Classifier is All You NeedSagar Vaze, Kai Han, Andrea Vedaldi, Andrew ZissermanICLR 2022 · 594 citations
Related papers
- Pedestrian Attribute Recognition: A New Benchmark Dataset and a Large Language Model Augmented FrameworkJiandong Jin, Xiao Wang, Qian Zhu, Haiyang Wang et al.AAAI 2025 · 19 citations
- Learning Disentangled Attribute Representations for Robust Pedestrian Attribute RecognitionJian Jia, Naiyu Gao, Fei He, Xiaotang Chen et al.AAAI 2022 · 50 citations
- Selective and Orthogonal Feature Activation for Pedestrian Attribute RecognitionJunyi Wu, Yan Huang, Min Gao, Yuzhen Niu et al.AAAI 2024 · 21 citations
- Joint Implicit and Explicit Language Learning for Pedestrian Attribute RecognitionYukang Zhang, Lei Tan, Yang Lu, Yan Yan et al.AAAI 2026 · 1 citation
- Enhanced Visual-Semantic Interaction with Tailored Prompts for Pedestrian Attribute RecognitionJunyi Wu, Yan Huang, Min Gao, Yuzhen Niu et al.CVPR 2025
