Interleaved Selective State Space Models for Efficient WiFi-Based 3D Multi-Person Pose Estimation
Quang-Anh N.D., Kok-Seng Wong
Abstract
WiFi-based human pose estimation offers privacy-preserving and occlusion-robust sensing, but current Transformer-based approaches suffer from quadratic complexity and lack explicit inductive biases for the structure of Channel State Information (CSI). We propose WiFi-Mamba, the first State Space Model (SSM) architecture for WiFi-based 3D multi-person pose estimation. Our approach introduces three key contributions: (1) a Dual-Stream Selective SSM that processes amplitude and phase through parallel pathways with cross-stream state coupling to respect their distinct physical properties, (2) Selective State Attention for pose query decoding with SSM-derived sequential context, and (3) Persistent SSM Memory for temporal consistency across frames without recurrent memory explosion. Extensive experiments on the Person-in-WiFi 3D dataset, covering both single-person and multi-person, demonstrate a 16-27% MPJPE reduction across varying numbers of persons while using only 4.4% of the baseline parameters (2.14M vs. 48.2M), achieving superior efficiency-accuracy trade-offs particularly beneficial for edge deployment in privacy-sensitive continuous monitoring scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a8fd46b9-6a4f-4637-8bae-95987b3377c7Builds on15
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 1,407 citations
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab et al.NeurIPS 2021 · 1,280 citations
Related papers
- PoseMamba: Monocular 3D Human Pose Estimation with Bidirectional Global-Local Spatio-Temporal State Space ModelYunlong Huang, Junshuo Liu, Ke Xian, Robert Caiming QiuAAAI 2025 · 15 citations
- Person-in-WiFi 3D: End-to-End Multi-Person 3D Pose Estimation with Wi-FiKangwei Yan, Fei Wang, Bo Qian, Han Ding et al.CVPR 2024 · 22 citations
- MV-SSM: Multi-View State Space Modeling for 3D Human Pose EstimationAviral Chharia, Wenbo Gou, Haoye DongCVPR 2025
- RFMamba: Frequency-Aware State Space Model for RF-Based Human-Centric PerceptionRui Zhang, Ruixu Geng, Yadong Li, Ruiyuan Song et al.ICLR 2025
- 3DET-Mamba: Causal Sequence Modelling for End-to-End 3D Object DetectionMingsheng Li, Jiakang Yuan, Sijin Chen, Lin Zhang et al.NeurIPS 2024 · 5 citations
