HumanBench: Towards General Human-Centric Perception with Projector Assisted Pretraining
Shixiang Tang, Cheng Chen, Qingsong Xie, Meilin Chen, Yizhou Wang, Yuanzheng Ci, Lei Bai, Feng Zhu, Haiyang Yang, Li Yi, Rui Zhao, Wanli Ouyang
Abstract
Human-centric perceptions include a variety of vision tasks, which have widespread industrial applications, including surveillance, autonomous driving, and the metaverse. It is desirable to have a general pretrain model for versatile human-centric downstream tasks. This paper forges ahead along this path from the aspects of both benchmark and pretraining methods. Specifically, we propose a HumanBench based on existing datasets to comprehensively evaluate on the common ground the generalization abilities of different pretraining methods on 19 datasets from 6 diverse downstream tasks, including person ReID, pose estimation, human parsing, pedestrian attribute recognition, pedestrian detection, and crowd counting. To learn both coarse-grained and fine-grained knowledge in human bodies, we further propose a Projector AssisTed Hierarchical pretraining method (PATH) to learn diverse knowledge at different granularity levels. Comprehensive evaluations on HumanBench show that our PATH achieves new state-of-the-art results on 17 downstream datasets and on-par results on the other 2 datasets. The code will be publicly at https://github.com/OpenGVLab/HumanBench .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- HAP: Structure-Aware Masked Image Modeling for Human-Centric PerceptionJunkun Yuan, Xinyu Zhang, Hao Zhou, Jian Wang et al.NeurIPS 2023 · 46 citations
- Cross-video Identity Correlating for Person Re-identification Pre-trainingJialong Zuo, Ying Nie, Hanyu Zhou, Huaxin Zhang et al.NeurIPS 2024 · 15 citations
- Structural Information Guided Multimodal Pre-training for Vehicle-Centric PerceptionXiao Wang, Wentao Wu, Chenglong Li, Zhicheng Zhao et al.AAAI 2024 · 10 citations
- CHASE: Learning Convex Hull Adaptive Shift for Skeleton-based Multi-Entity Action RecognitionYuhang Wen, Mengyuan Liu, Songtao Wu, Beichen DingNeurIPS 2024 · 7 citations
- RELI11D: A Comprehensive Multimodal Human Motion Dataset and MethodMing Yan, Yan Zhang, Shuqiang Cai, Shuqi Fan et al.CVPR 2024 · 5 citations
Builds on35
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
Related papers
- UniHCP: A Unified Model for Human-Centric PerceptionsYuanzheng Ci, Yizhou Wang, Meilin Chen, Shixiang Tang et al.CVPR 2023
- Cross-View and Cross-Pose Completion for 3D Human UnderstandingMatthieu Armando, Salma Galaaoui, Fabien Baradel, Thomas Lucas et al.CVPR 2024
- Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language ModelsYuansen Liu, Haiming Tang, Jinlong Peng, Jiangning Zhang et al.ICLR 2026 · 5 citations
- Generalizable Pedestrian Detection: The Elephant in the RoomIrtiza Hasan, Shengcai Liao, Jinpeng Li, Saad Ullah Akram et al.CVPR 2021
- MotionBERT: A Unified Perspective on Learning Human Motion RepresentationsWentao Zhu, Xiaoxuan Ma, Zhaoyang Liu, Libin Liu et al.ICCV 2023 · 322 citations
