Two-stage Co-segmentation Network Based on Discriminative Representation for Recovering Human Mesh from Videos
Boyang Zhang, Kehua Ma, Suping Wu, Zhixiang Yuan
摘要
Recovering 3D human mesh from videos has recently made significant progress. However, most of the existing methods focus on the temporal consistency of videos, while ignoring the spatial representation in complex scenes, thus failing to recover a reasonable and smooth human mesh sequence under extreme illumination and chaotic backgrounds. To alleviate this problem, we propose a two-stage co-segmentation network based on discriminative representationfor recovering human body meshes from videos. Specifically, the first stage of the network segments the video spatial domain to spotlight spatially fine-grained information, and then learns and enhances the intra-frame discriminative representation through a dual-excitation mechanism and a frequency domain enhancement module, while sup-pressing irrelevant information (e.g., background). The second stage focuses on temporal context by segmenting the video temporal domain, and models inter-frame discriminative representation via a dynamic integration strategy. Further, to efficiently generate reasonable human discriminative actions, we carefully elaborate a landmark anchor area loss to constrain the variation of the human motion area. Extensive experimental results on large publicly available datasets indicate superiority in comparison with most state-of-the-art. The Code will be made public.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- PhysPT: Physics-aware Pretrained Transformer for Estimating Human Dynamics from Monocular VideosYufei Zhang, Jeffrey O. Kephart, Zijun Cui, Qiang JiCVPR 2024 · 被引用 14 次
- ARTS: Semi-Analytical Regressor using Disentangled Skeletal Representations for Human Mesh Recovery from VideosTao Tang, Hong Liu, Yingxuan You, Ti Wang 等ACM MM 2024 · 被引用 2 次
它引用的顶会 Paper7
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 被引用 1,139 次
- Human Mesh Recovery From Monocular Images via a Skeleton-Disentangled RepresentationYu Sun, Yun Ye, Wu Liu, Wenpeng Gao 等ICCV 2019 · 被引用 196 次
- Capturing Humans in Motion: Temporal-Attentive 3D Human Pose and Shape Estimation from Monocular VideoWen-Li Wei, Jen-Chun Lin, Tyng-Luh Liu, Hong-Yuan Mark LiaoCVPR 2022 · 被引用 117 次
- EventHPE: Event-based 3D Human Pose and Shape EstimationShihao Zou, Chuan Guo, Xinxin Zuo, Sen Wang 等ICCV 2021 · 被引用 77 次
- DC-GNet: Deep Mesh Relation Capturing Graph Convolution Network for 3D Human Shape ReconstructionShihao Zhou, Mengxi Jiang, Shanshan Cai, Yunqi LeiACM MM 2021 · 被引用 14 次
相关 Paper
- Vid2Avatar: 3D Avatar Reconstruction from Videos in the Wild via Self-supervised Scene DecompositionChen Guo, Tianjian Jiang, Xu Chen, Jie Song 等CVPR 2023
- SelfRecon: Self Reconstruction Your Digital Avatar from Monocular VideoBoyi Jiang, Yang Hong, Hujun Bao, Juyong ZhangCVPR 2022 · 被引用 142 次
- Co-Evolution of Pose and Mesh for 3D Human Body Estimation from VideoYingxuan You, Hong Liu, Ti Wang, Wenhao Li 等ICCV 2023 · 被引用 35 次
- Clip Fusion with Bi-level Optimization for Human Mesh Reconstruction from Monocular VideosPeng Wu, Xiankai Lu, Jianbing Shen, Yilong YinACM MM 2023 · 被引用 18 次
- Cyclic Test-Time Adaptation on Monocular Video for 3D Human Mesh ReconstructionHyeongjin Nam, Daniel Sungho Jung, Yeonguk Oh, Kyoung Mu LeeICCV 2023 · 被引用 28 次
