JOTR: 3D Joint Contrastive Learning with Transformers for Occluded Human Mesh Recovery
Jiahao Li, Zongxin Yang, Xiaohan Wang, Jianxin Ma, Chang Zhou, Yi Yang
Abstract
In this study, we focus on the problem of 3D human mesh recovery from a single image under obscured conditions. Most state-of-the-art methods aim to improve 2D alignment technologies, such as spatial averaging and 2D joint sampling. However, they tend to neglect the crucial aspect of 3D alignment by improving 3D representations. Furthermore, recent methods struggle to separate the target human from occlusion or background in crowded scenes as they optimize the 3D space of target human with 3D joint coordinates as local supervision. To address these issues, a desirable method would involve a framework for fusing 2D and 3D features and a strategy for optimizing the 3D space globally. Therefore, this paper presents 3D JOint contrastive learning with TRansformers (JOTR) framework for handling occluded 3D human mesh recovery. Our method includes an encoderdecoder transformer architecture to fuse 2D and 3D representations for achieving 2D&3D aligned results in a coarseto-fine manner and a novel 3D joint contrastive learning approach for adding explicitly global supervision for the 3D † Jiahao Li worked on this at his Alibaba internship. ‡ Yi Yang is the corresponding author. feature space. The contrastive learning approach includes two contrastive losses: joint-to-joint contrast for enhancing the similarity of semantically similar voxels (i.e., human joints), and joint-to-non-joint contrast for ensuring discrimination from others (e.g., occlusions and background). Qualitative and quantitative analyses demonstrate that our method outperforms state-of-the-art competitors on both occlusionspecific and standard benchmarks, significantly improving the reconstruction of occluded humans. Code is available at https://github.com/xljh0520/JOTR .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bf4b23c0-099a-421b-899c-b6313f63c57eCited by top-tier papers6
- SIFU: Side-view Conditioned Implicit Function for Real-world Usable Clothed Human ReconstructionZechuan Zhang, Zongxin Yang, Yi YangCVPR 2024 · 44 citations
- RoHM: Robust Human Motion Reconstruction via DiffusionSiwei Zhang, Bharat Lal Bhatnagar, Yuanlu Xu, Alexander Winkler et al.CVPR 2024 · 11 citations
- DPMesh: Exploiting Diffusion Prior for Occluded Human Mesh RecoveryYixuan Zhu, Ao Li, Yansong Tang, Wenliang Zhao et al.CVPR 2024 · 10 citations
- ScoreHypo: Probabilistic Human Mesh Estimation with Hypothesis ScoringYuan Xu, Xiaoxuan Ma, Jiajun Su, Wentao Zhu et al.CVPR 2024 · 6 citations
- Occluded Human Body Capture with Frequency Domain Denoising PriorBuzhen Huang, Chongyang Xu, Wentao Tang, Yuan Shu et al.CVPR 2026
Builds on42
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
Related papers
- Instance-Aware Contrastive Learning for Occluded Human Mesh ReconstructionMi-Gyeong Gwon, Gi-Mun Um, Won-Sik Cheong, Wonjun KimCVPR 2024
- OCR-Pose: Occlusion-aware Contrastive Representation for Unsupervised 3D Human Pose EstimationJunjie Wang, Zhenbo Yu, Zhengyan Tong, Hang Wang et al.ACM MM 2022 · 12 citations
- LiftedCL: Lifting Contrastive Learning for Human-Centric PerceptionZiwei Chen, Qiang Li, Xiaofeng Wang, Wankou YangICLR 2023
- Learning Human Mesh Recovery in 3D ScenesZehong Shen, Zhi Cen, Sida Peng, Qing Shuai et al.CVPR 2023
- 3D Human Mesh Reconstruction by Learning to Sample Joint Adaptive Tokens for TransformersYouze Xue, Jiansheng Chen, Yudong Zhang, Cheng Yu et al.ACM MM 2022 · 10 citations
