Clip Fusion with Bi-level Optimization for Human Mesh Reconstruction from Monocular Videos
Peng Wu, Xiankai Lu, Jianbing Shen, Yilong Yin
Abstract
Human mesh reconstruction (HMR) from monocular video is the key step to many mixed reality and robotic applications. Although existing methods show promising results by capturing frames' temporal information, these methods predict human mesh with the design of implicit temporal learning modules in a sequence to frame manner. To mine more temporal information from the video, we present a bi-level clip inference network for HMR, which leverages both local motion and global context explicitly for dense 3D reconstruction. Specifically, we propose a novel bi-level temporal fusion strategy that takes both neighboring and long-range relations into consideration. In addition, different from traditional frame-wise operation, we investigate an alternative perspective by treating video-based HMR as clip-wise inference. We evaluate the proposed method on multiple datasets (3DPW, Human3.6M, and MPI-INF-3DHP) quantitatively and qualitatively, demonstrating a significant improvement over existing methods (in terms of PA-MPJPE, ACC-Error etc). Furthermore, we extend the proposed method on more challenging Multiple Shots HMR task to demonstrate its generalizability. Some visual demos can be seen https://github.com/bicf0/bicf_demo.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers3
- ARTS: Semi-Analytical Regressor using Disentangled Skeletal Representations for Human Mesh Recovery from VideosTao Tang, Hong Liu, Yingxuan You, Ti Wang et al.ACM MM 2024 · 2 citations
- Towards Practical Human Motion Prediction with LiDAR Point CloudsXiao Han, Yiming Ren, Yichen Yao, Yujing Sun et al.ACM MM 2024 · 2 citations
- HumanMM: Global Human Motion Recovery from Multi-shot VideosYuhong Zhang, Guanlin Wu, Ling-Hao Chen, Zhuokai Zhao et al.CVPR 2025
Related papers
- Capturing Humans in Motion: Temporal-Attentive 3D Human Pose and Shape Estimation from Monocular VideoWen-Li Wei, Jen-Chun Lin, Tyng-Luh Liu, Hong-Yuan Mark LiaoCVPR 2022 · 117 citations
- Co-Evolution of Pose and Mesh for 3D Human Body Estimation from VideoYingxuan You, Hong Liu, Ti Wang, Wenhao Li et al.ICCV 2023 · 35 citations
- Two-stage Co-segmentation Network Based on Discriminative Representation for Recovering Human Mesh from VideosBoyang Zhang, Kehua Ma, Suping Wu, Zhixiang YuanCVPR 2023
- MHFormer: Multi-Hypothesis Transformer for 3D Human Pose EstimationWenhao Li, Hong Liu, Hao Tang, Pichao Wang et al.CVPR 2022 · 403 citations
- GLAMR: Global Occlusion-Aware Human Mesh Recovery with Dynamic CamerasYe Yuan, Umar Iqbal, Pavlo Molchanov, Kris Kitani et al.CVPR 2022 · 111 citations
