MV-SSM: Multi-View State Space Modeling for 3D Human Pose Estimation
Aviral Chharia, Wenbo Gou, Haoye Dong
摘要
While significant progress has been made in single-view 3D human pose estimation, multi-view 3D human pose estimation remains challenging, particularly in terms of generalizing to new camera configurations. Existing attention-based transformers often struggle to accurately model the spatial arrangement of keypoints, especially in occluded scenarios. Additionally, they tend to overfit specific camera arrangements and visual scenes from training data, resulting in substantial performance drops in new settings. In this study, we introduce a novel Multi-View State Space Modeling framework, named MV-SSM, for robustly estimating 3D human keypoints. We explicitly model the joint spatial sequence at two distinct levels: the feature level from multi-view images and the person keypoint level. We propose a Projective State Space (PSS) block to learn a generalized representation of joint spatial arrangements using state space modeling. Moreover, we modify Mamba's traditional scanning into an effective Grid Token-guided Bidirectional Scanning (GTBS), which is integral to the PSS block. Multiple experiments demonstrate that MV-SSM achieves strong generalization, outperforming state-of-the-art methods: +10.8 on AP 25 (+24% ↑) on the challenging three-camera setting in CMU Panoptic, +7.0 on AP 25 (+13% ↑) on varying camera arrangements, and +15.3 PCP (+38% ↑) on Campus A1 in cross-dataset evaluations. Project Website: https: //aviralchharia.github.io/MV-SSM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Differentially Private 2D Human Pose EstimationKaushik Bhargav Sivangi, Paul Henderson, Fani DeligianniCVPR 2026 · 被引用 1 次
- Person Parametric Physics-informed Representation for mmWave-based Human Pose EstimationShuntian Zheng, Jiaqi Li, Guangming Wang, Minzhe Ni 等UbiComp 2026 · 被引用 1 次
- PS-Mamba: Spatial-Temporal Graph Mamba for Pose Sequence RefinementHaoye Dong, Gim Hee LeeICCV 2025
- DisPOSE: Projected Polystochastic Diffusion for Self-Supervised Multi-View 3D Human Pose EstimationTony Danjun Wang, Tolga Birdal, Nassir Navab, Lennart BastianICML 2026
- Multi-view Consistent 3D Gaussian Head Avatars 'without' Multi-view GenerationAviral Chharia, Fernando De la TorreCVPR 2026
它引用的顶会 Paper16
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang 等ICML 2024 · 被引用 1,725 次
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab 等NeurIPS 2021 · 被引用 1,280 次
- Learnable Triangulation of Human PoseKarim Iskakov, Egor Burkov, Victor S. Lempitsky, Yury MalkovICCV 2019 · 被引用 419 次
相关 Paper
- PoseMamba: Monocular 3D Human Pose Estimation with Bidirectional Global-Local Spatio-Temporal State Space ModelYunlong Huang, Junshuo Liu, Ke Xian, Robert Caiming QiuAAAI 2025 · 被引用 15 次
- High-Resolution Spatiotemporal Modeling with Global-Local State Space Models for Video-Based Human Pose EstimationRunyang Feng, Hyung Jin Chang, Tze Ho Elden Tse, Boeun Kim 等ICCV 2025 · 被引用 2 次
- Direct Multi-view Multi-person 3D Pose EstimationTao Wang, Jianfeng Zhang, Yujun Cai, Shuicheng Yan 等NeurIPS 2021 · 被引用 147 次
- Interleaved Selective State Space Models for Efficient WiFi-Based 3D Multi-Person Pose EstimationQuang-Anh N.D., Kok-Seng WongICML 2026
- HandMCM: Multi-modal Point Cloud-based Correspondence State Space Model for 3D Hand Pose EstimationWencan Cheng, Gim Hee LeeAAAI 2026
