Lune

ACM MM2025顶会

DSP: Dense-Sparse Parallel Networks for Self-supervised 3D Multi-person Pose Estimation from Multiple Views

Yang Liu, Zhiyong Zhang

2025年份
1顶会引用

摘要

Recent advances in multi-view 3D multi-person pose estimation have led to significant progress. However, several critical challenges remain, including the limited extraction and integration of multi-domain information, as well as the high annotation costs associated with 3D data in multi-person scenarios. These issues hinder the broader applicability of current methods in complex computer vision tasks. In this paper, we propose a Dense-Sparse Parallel Networks (DSP) framework that jointly leverages spatial, temporal, and frequency-domain information through an adaptive geo-consistency self-supervised strategy. Specifically, we design a multi-view spatial feature extraction module that captures cross-view spatial distributions from dense multi-view feature maps. In parallel, we employ a local-global temporal attention module and a frequency-aware attention module to extract dynamic temporal patterns and localized frequency-domain features from sparse keypoint data. Furthermore, a multi-domain parallel fusion module is introduced to effectively integrate features across all domains, enabling accurate multi-person 3D pose regression. To enhance self-supervised learning, we employ a dynamic view selector guided by reinforcement learning, which reduces the impact of inaccurate pre-trained 2D poses. Experimental results on three benchmark datasets (i.e., CMU Panoptic, Campus, and Shelf) demonstrate that the proposed DSP framework achieves robust and accurate performance, as evidenced by comparisons with other state-of-the-art methods.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖