Two-stage Co-segmentation Network Based on Discriminative Representation for Recovering Human Mesh from Videos
Boyang Zhang, Kehua Ma, Suping Wu, Zhixiang Yuan
Abstract
Recovering 3D human mesh from videos has recently made significant progress. However, most of the existing methods focus on the temporal consistency of videos, while ignoring the spatial representation in complex scenes, thus failing to recover a reasonable and smooth human mesh sequence under extreme illumination and chaotic backgrounds. To alleviate this problem, we propose a two-stage co-segmentation network based on discriminative representationfor recovering human body meshes from videos. Specifically, the first stage of the network segments the video spatial domain to spotlight spatially fine-grained information, and then learns and enhances the intra-frame discriminative representation through a dual-excitation mechanism and a frequency domain enhancement module, while sup-pressing irrelevant information (e.g., background). The second stage focuses on temporal context by segmenting the video temporal domain, and models inter-frame discriminative representation via a dynamic integration strategy. Further, to efficiently generate reasonable human discriminative actions, we carefully elaborate a landmark anchor area loss to constrain the variation of the human motion area. Extensive experimental results on large publicly available datasets indicate superiority in comparison with most state-of-the-art. The Code will be made public.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- PhysPT: Physics-aware Pretrained Transformer for Estimating Human Dynamics from Monocular VideosYufei Zhang, Jeffrey O. Kephart, Zijun Cui, Qiang JiCVPR 2024 · 14 citations
- ARTS: Semi-Analytical Regressor using Disentangled Skeletal Representations for Human Mesh Recovery from VideosTao Tang, Hong Liu, Yingxuan You, Ti Wang et al.ACM MM 2024 · 2 citations
Builds on7
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- Human Mesh Recovery From Monocular Images via a Skeleton-Disentangled RepresentationYu Sun, Yun Ye, Wu Liu, Wenpeng Gao et al.ICCV 2019 · 196 citations
- Capturing Humans in Motion: Temporal-Attentive 3D Human Pose and Shape Estimation from Monocular VideoWen-Li Wei, Jen-Chun Lin, Tyng-Luh Liu, Hong-Yuan Mark LiaoCVPR 2022 · 117 citations
- EventHPE: Event-based 3D Human Pose and Shape EstimationShihao Zou, Chuan Guo, Xinxin Zuo, Sen Wang et al.ICCV 2021 · 77 citations
- DC-GNet: Deep Mesh Relation Capturing Graph Convolution Network for 3D Human Shape ReconstructionShihao Zhou, Mengxi Jiang, Shanshan Cai, Yunqi LeiACM MM 2021 · 14 citations
Related papers
- Vid2Avatar: 3D Avatar Reconstruction from Videos in the Wild via Self-supervised Scene DecompositionChen Guo, Tianjian Jiang, Xu Chen, Jie Song et al.CVPR 2023
- SelfRecon: Self Reconstruction Your Digital Avatar from Monocular VideoBoyi Jiang, Yang Hong, Hujun Bao, Juyong ZhangCVPR 2022 · 142 citations
- Co-Evolution of Pose and Mesh for 3D Human Body Estimation from VideoYingxuan You, Hong Liu, Ti Wang, Wenhao Li et al.ICCV 2023 · 35 citations
- Clip Fusion with Bi-level Optimization for Human Mesh Reconstruction from Monocular VideosPeng Wu, Xiankai Lu, Jianbing Shen, Yilong YinACM MM 2023 · 18 citations
- Cyclic Test-Time Adaptation on Monocular Video for 3D Human Mesh ReconstructionHyeongjin Nam, Daniel Sungho Jung, Yeonguk Oh, Kyoung Mu LeeICCV 2023 · 28 citations
