Joint Implicit and Explicit Language Learning for Pedestrian Attribute Recognition
Yukang Zhang, Lei Tan, Yang Lu, Yan Yan, Hanzi Wang
摘要
Pedestrian attribute recognition (PAR) has received increasing attention due to its wide application in video surveillance and pedestrian analysis. Some text-enhanced methods tackle this task by facilitating interactive learning between the attributes and the visual images. However, these generic languages fail to uniquely describe different pedestrian images, missing individual characteristics. In this paper, we propose a Joint Implicit and Explicit Language Guidance Enhancement Learning (JGEL) method, which converts each pedestrian image into a language description with dual language learning to effectively learn enhanced attribute information. Specifically, we first propose an Implicit Language Guidance Learning (ILGL) stream. It projects visual image features into the text embedding space to generate pseudo-word tokens, implicitly modeling image attributes and providing personalized descriptions. Moreover, we propose an Explicit Attribute Enhancement Learning (EAEL) stream to guide the generated pseudo-word tokens obtained by ILGL explicitly aligned with pedestrian attributes, which can effectively align the pseudoword tokens with the attribute concepts in the text embedding space. Extensive experiments show that JGEL has significant advantages in improving the performance of PAR and the challenging zero-shot PAR task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- SimVLM: Simple Visual Language Model Pretraining with Weak SupervisionZirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai 等ICLR 2022 · 被引用 950 次
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual InversionRinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik 等ICLR 2023 · 被引用 464 次
- CLIP-ReID: Exploiting Vision-Language Model for Image Re-identification without Concrete Text LabelsSiyuan Li, Li Sun, Qingli LiAAAI 2023 · 被引用 355 次
相关 Paper
- Enhanced Visual-Semantic Interaction with Tailored Prompts for Pedestrian Attribute RecognitionJunyi Wu, Yan Huang, Min Gao, Yuzhen Niu 等CVPR 2025
- POAR: Towards Open Vocabulary Pedestrian Attribute RecognitionYue Zhang, Suchen Wang, Shichao Kan, Zhenyu Weng 等ACM MM 2023 · 被引用 11 次
- Improving Pedestrian Attribute Recognition With Weakly-Supervised Multi-Scale Attribute-Specific LocalizationChufeng Tang, Lu Sheng, Zhaoxiang Zhang, Xiaolin HuICCV 2019 · 被引用 153 次
- Relation-Aware Pedestrian Attribute Recognition with Graph Convolutional NetworksZichang Tan, Yang Yang, Jun Wan, Guodong Guo 等AAAI 2020 · 被引用 103 次
- Cross-modal Co-occurrence Attributes Alignments for Person Search by LanguageKai Niu, Linjiang Huang, Yan Huang, Peng Wang 等ACM MM 2022 · 被引用 35 次
