HAP: Structure-Aware Masked Image Modeling for Human-Centric Perception
Junkun Yuan, Xinyu Zhang, Hao Zhou, Jian Wang, Zhongwei Qiu, Zhiyin Shao, Shaofeng Zhang, Sifan Long, Kun Kuang, Kun Yao, Junyu Han, Errui Ding
摘要
Model pre-training is essential in human-centric perception. In this paper, we first introduce masked image modeling (MIM) as a pre-training approach for this task. Upon revisiting the MIM training strategy, we reveal that human structure priors offer significant potential. Motivated by this insight, we further incorporate an intuitive human structure prior - human parts - into pre-training. Specifically, we employ this prior to guide the mask sampling process. Image patches, corresponding to human part regions, have high priority to be masked out. This encourages the model to concentrate more on body structure information during pre-training, yielding substantial benefits across a range of human-centric perception tasks. To further capture human characteristics, we propose a structure-invariant alignment loss that enforces different masked views, guided by the human part prior, to be closely aligned for the same image. We term the entire method as HAP. HAP simply uses a plain ViT as the encoder yet establishes new state-of-the-art performance on 11 human-centric benchmarks, and on-par result on one dataset. For example, HAP achieves 78.1% mAP on MSMT17 for person re-identification, 86.54% mA on PA-100K for pedestrian attribute recognition, 78.2% AP on MS COCO for 2D pose estimation, and 56.0 PA-MPJPE on 3DPW for 3D pose and shape estimation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image SynthesisYanzuo Lu, Manlin Zhang, Andy J. Ma, Xiaohua Xie 等CVPR 2024 · 被引用 26 次
- Follow-Your-Preference: Towards Preference-Aligned Image InpaintingYutao Shen, Junkun Yuan, Toru Aonishi, Hideki Nakayama 等ICLR 2026 · 被引用 21 次
- Pedestrian Attribute Recognition: A New Benchmark Dataset and a Large Language Model Augmented FrameworkJiandong Jin, Xiao Wang, Qian Zhu, Haiyang Wang 等AAAI 2025 · 被引用 19 次
- CHASE: Learning Convex Hull Adaptive Shift for Skeleton-based Multi-Entity Action RecognitionYuhang Wen, Mengyuan Liu, Songtao Wu, Beichen DingNeurIPS 2024 · 被引用 7 次
- Asynchronous Denoising Diffusion Models for Aligning Text-to-Image GenerationZijing Hu, Yunze Tong, Fengda Zhang, Junkun Yuan 等ICLR 2026 · 被引用 3 次
它引用的顶会 Paper42
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 被引用 2,340 次
相关 Paper
- Cross-View and Cross-Pose Completion for 3D Human UnderstandingMatthieu Armando, Salma Galaaoui, Fabien Baradel, Thomas Lucas 等CVPR 2024
- UniHCP: A Unified Model for Human-Centric PerceptionsYuanzheng Ci, Yizhou Wang, Meilin Chen, Shixiang Tang 等CVPR 2023
- LiftedCL: Lifting Contrastive Learning for Human-Centric PerceptionZiwei Chen, Qiang Li, Xiaofeng Wang, Wankou YangICLR 2023
- Versatile Multi-Modal Pre-Training for Human-Centric PerceptionFangzhou Hong, Liang Pan, Zhongang Cai, Ziwei LiuCVPR 2022 · 被引用 15 次
- Integrally Pre-Trained Transformer Pyramid NetworksYunjie Tian, Lingxi Xie, Zhaozhi Wang, Longhui Wei 等CVPR 2023
