Cross-View and Cross-Pose Completion for 3D Human Understanding
Matthieu Armando, Salma Galaaoui, Fabien Baradel, Thomas Lucas, Vincent Leroy, Romain Brégier, Philippe Weinzaepfel, Grégory Rogez
Abstract
Abstract Human perception and understanding is a major domain of computer vision which, like many other vision subdomains recently, stands to gain from the use of large models pre-trained on large datasets. We hypothesize that the most common pre-training strategy of relying on general purpose, object-centric image datasets such as ImageNet, is limited by an important domain shift. On the other hand, collecting domain-specific ground truth such as 2D or 3D labels does not scale well. Therefore, we propose a pre-training approach based on self-supervised learning that works on human-centric data using only images. Our method uses pairs of images of humans: the first is partially masked and the model is trained to reconstruct the masked parts given the visible ones and a second image. It relies on both stereoscopic (cross-view) pairs, and temporal (cross-pose) pairs taken from videos, in order to learn priors about 3D as well as human motion. We pre-train a model for body-centric tasks and one for hand-centric tasks. With a generic transformer architecture, these models outperform existing self-supervised pre-training methods on a wide set of human-centric downstream tasks, and obtain state-of-the-art performance for instance when fine-tuning for model-based and model-free human mesh recovery.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7207bc6d-80bd-44e1-a7fa-650208bf25caCited by top-tier papers2
- EAGLE: Efficient Adaptive Geometry-based Learning in Cross-view UnderstandingThanh-Dat Truong, Utsav Prabhu, Dongyi Wang, Bhiksha Raj et al.NeurIPS 2024 · 7 citations
- MEGA: Masked Generative Autoencoder for Human Mesh RecoveryGuénolé Fiche, Simon Leglaive, Xavier Alameda-Pineda, Francesc Moreno-NoguerCVPR 2025
Builds on37
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
Related papers
- CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View CompletionPhilippe Weinzaepfel, Vincent Leroy, Thomas Lucas, Romain Brégier et al.NeurIPS 2022 · 189 citations
- LiftedCL: Lifting Contrastive Learning for Human-Centric PerceptionZiwei Chen, Qiang Li, Xiaofeng Wang, Wankou YangICLR 2023
- HAP: Structure-Aware Masked Image Modeling for Human-Centric PerceptionJunkun Yuan, Xinyu Zhang, Hao Zhou, Jian Wang et al.NeurIPS 2023 · 46 citations
- Mask3D: Pretraining 2D Vision Transformers by Learning Masked 3D PriorsJi Hou, Xiaoliang Dai, Zijian He, Angela Dai et al.CVPR 2023
- Self-supervised Transfer Learning for Hand Mesh Recovery from Binocular ImagesZheng Chen, Sihan Wang, Yi Sun, Xiaohong MaICCV 2021 · 7 citations
