POAR: Towards Open Vocabulary Pedestrian Attribute Recognition
Yue Zhang, Suchen Wang, Shichao Kan, Zhenyu Weng, Yigang Cen, Yap-Peng Tan
摘要
Pedestrian attribute recognition (PAR) aims to predict the attributes of a target pedestrian. Recent methods often address the PAR problem by training a multi-label classifier with predefined attribute classes, but they can hardly exhaust all possible pedestrian attributes in the real world. To tackle this problem, we propose a novel Pedestrian Open-Attribute Recognition (POAR) approach by formulating the problem as a task of image-text search. Our approach employs a Transformer-based Encoder with a Masking Strategy (TEMS) to focus on the attributes of specific pedestrian parts (e.g., head, upper body, lower body, feet, etc.), and introduces a set of attribute tokens to encode the corresponding attributes into visual embeddings. Each attribute category is described as a natural language sentence and encoded by the text encoder. Then, we compute the similarity between the visual and text embeddings to find the best attribute descriptions for the input images. To handle multiple attributes of a single pedestrian, we propose a Many-To-Many Contrastive (MTMC) loss with masked tokens. In addition, we propose a Grouped Knowledge Distillation (GKD) method to minimize the disparity between visual embeddings and unseen attribute text embeddings. We evaluate our proposed method on three PAR datasets with an open-attribute setting. The results demonstrate the effectiveness of our method as a strong baseline for the POAR task. Our code is available at https://github.com/IvyYZ/POAR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
- TransReID: Transformer-based Object Re-IdentificationShuting He, Hao Luo, Pichao Wang, Fan Wang 等ICCV 2021 · 被引用 1,172 次
- Open-Set Recognition: A Good Closed-Set Classifier is All You NeedSagar Vaze, Kai Han, Andrea Vedaldi, Andrew ZissermanICLR 2022 · 被引用 594 次
相关 Paper
- Pedestrian Attribute Recognition: A New Benchmark Dataset and a Large Language Model Augmented FrameworkJiandong Jin, Xiao Wang, Qian Zhu, Haiyang Wang 等AAAI 2025 · 被引用 19 次
- Learning Disentangled Attribute Representations for Robust Pedestrian Attribute RecognitionJian Jia, Naiyu Gao, Fei He, Xiaotang Chen 等AAAI 2022 · 被引用 50 次
- Selective and Orthogonal Feature Activation for Pedestrian Attribute RecognitionJunyi Wu, Yan Huang, Min Gao, Yuzhen Niu 等AAAI 2024 · 被引用 21 次
- Joint Implicit and Explicit Language Learning for Pedestrian Attribute RecognitionYukang Zhang, Lei Tan, Yang Lu, Yan Yan 等AAAI 2026 · 被引用 1 次
- Enhanced Visual-Semantic Interaction with Tailored Prompts for Pedestrian Attribute RecognitionJunyi Wu, Yan Huang, Min Gao, Yuzhen Niu 等CVPR 2025
