From Sparse to Dense: Spatio-Temporal Fusion for Multi-View 3D Human Pose Estimation with DenseWarper
Ling Li, Changjie Chen, Yuyan Wang, Jiaqing Lyu, Kenglun Chang, Yiyun Chen, Zhidong Deng
Abstract
In multi-view 3D human pose estimation, models typically rely on images captured simultaneously from different camera views to predict a pose at a specific moment. While providing accurate spatial information, this traditional approach often overlooks the rich temporal dependencies between adjacent frames. We propose a novel 3D human pose estimation input method: the sparse interleaved input to address this. This method leverages images captured from different camera views at various time points (e.g., View 1 at time and View 2 at time ), allowing our model to capture rich spatio-temporal information and effectively boost performance. More importantly, this approach offers two key advantages: First, it can theoretically increase the output pose frame rate by N times with N cameras, thereby breaking through single-view frame rate limitations and enhancing the temporal resolution of the production. Second, using a sparse subset of available frames, our method can reduce data redundancy and simultaneously achieve better performance. We introduce the DenseWarper model, which leverages epipolar geometry for efficient spatio-temporal heatmap exchange. We conducted extensive experiments on the Human3.6M and MPI-INF-3DHP datasets. Results demonstrate that our method, utilizing only sparse interleaved images as input, outperforms traditional dense multi-view input approaches and achieves state-of-the-art performance. The source code for this work is available at: https://github.com/lingli1724/DenseWarper-ICLR2026
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9a613c68-03b3-48eb-b694-65edf4facda2Builds on36
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- 3D Human Pose Estimation with Spatial and Temporal TransformersCe Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang et al.ICCV 2021 · 648 citations
- Exploiting Spatial-Temporal Relationships for 3D Pose Estimation via Graph Convolutional NetworksYujun Cai, Liuhao Ge, Jun Liu, Jianfei Cai et al.ICCV 2019 · 504 citations
- Learnable Triangulation of Human PoseKarim Iskakov, Egor Burkov, Victor S. Lempitsky, Yury MalkovICCV 2019 · 419 citations
- MHFormer: Multi-Hypothesis Transformer for 3D Human Pose EstimationWenhao Li, Hong Liu, Hao Tang, Pichao Wang et al.CVPR 2022 · 403 citations
Related papers
- Deep Semantic Graph Transformer for Multi-View 3D Human Pose EstimationLijun Zhang, Kangkang Zhou, Feng Lu, Xiang-Dong Zhou et al.AAAI 2024 · 14 citations
- Cross-View Tracking for Multi-Human 3D Pose Estimation at Over 100 FPSLong Chen, Haizhou Ai, Rui Chen, Zijie Zhuang et al.CVPR 2020
- SVTformer: Spatial-View-Temporal Transformer for Multi-View 3D Human Pose EstimationWanruo Zhang, Mengyuan Liu, Hong Liu, Wenhao LiAAAI 2025 · 4 citations
- Efficient Hierarchical Multi-view Fusion Transformer for 3D Human Pose EstimationKangkang Zhou, Lijun Zhang, Feng Lu, Xiang-Dong Zhou et al.ACM MM 2023 · 17 citations
- Multi-View 3D Human Pose Estimation with Weakly Synchronized ImagesLing Li, Ruiwen Gu, Chongyang Wang, Junliang Xing et al.AAAI 2025 · 3 citations
