Skeleton MixFormer: Multivariate Topology Representation for Skeleton-based Action Recognition
Wentian Xin, Qiguang Miao, Yi Liu, Ruyi Liu, Chi-Man Pun, Cheng Shi
摘要
Vision Transformer, which performs well in various vision tasks, encounters a bottleneck in skeleton-based action recognition and falls short of advanced GCN-based methods. The root cause is that the current skeleton transformer depends on the self-attention mechanism of the complete channel of the global joint, ignoring the highly discriminative differential correlation within the channel, so it is challenging to learn the expression of the multivariate topology dynamically. To tackle this, we present Skeleton MixFormer, an innovative spatio-temporal architecture to effectively represent the physical correlations and temporal interactivity of the compact skeleton data. Two essential components make up the proposed framework: 1) Spatial MixFormer. The channel-grouping and mix-attention are utilized to calculate the dynamic multivariate topological relationships. Compared with the full-channel self-attention method, Spatial MixFormer better highlights the channel groups' discriminative differences and the joint adjacency's interpretable learning. 2) Temporal MixFormer, which consists of Multiscale Convolution, Temporal Transformer and Sequential Holding Module. The multivariate temporal models ensure the richness of global difference expression and realize the discrimination of crucial intervals in the sequence, thereby enabling more effective learning of long and short-term dependencies in actions. Our Skeleton MixFormer demonstrates state-of-the-art (SOTA) performance across seven different settings on four standard datasets, namely NTU-60, NTU-120, NW-UCLA, and UAV-Human. Related code will be available on https://github.com/ElricXin/Skeleton-MixFormer.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper6
- LLMs are Good Action RecognizersHaoxuan Qu, Yujun Cai, Jun LiuCVPR 2024 · 被引用 37 次
- Multi-Modality Co-Learning for Efficient Skeleton-based Action RecognitionJinfu Liu, Chen Chen, Mengyuan LiuACM MM 2024 · 被引用 27 次
- Frequency Guidance Matters: Skeletal Action Recognition by Frequency-Aware Mixed TransformerWenhan Wu, Ce Zheng, Zihao Yang, Chen Chen 等ACM MM 2024 · 被引用 16 次
- Frequency-Semantic Enhanced Variational Autoencoder for Zero-Shot Skeleton-Based Action RecognitionWenhan Wu, Zhishuai Guo, Chen Chen, Hongfei Xue 等ICCV 2025 · 被引用 4 次
- SkeletonContext: Skeleton-side Context Prompt Learning for Zero-Shot Skeleton-based Action RecognitionNing Wang, Tieyue Wu, Naeha Sharif, Farid Boussaïd 等CVPR 2026 · 被引用 3 次
相关 Paper
- Spatio-Temporal Fusion for Human Action Recognition via Joint Trajectory GraphYaolin Zheng, Hongbo Huang, Xiuying Wang, Xiaoxu Yan 等AAAI 2024 · 被引用 22 次
- Topology-Aware Convolutional Neural Network for Efficient Skeleton-Based Action RecognitionKailin Xu, Fanfan Ye, Qiaoyong Zhong, Di XieAAAI 2022 · 被引用 168 次
- Multi-Scale Spatial Temporal Graph Convolutional Network for Skeleton-Based Action RecognitionZhan Chen, Sicheng Li, Bing Yang, Qinghan Li 等AAAI 2021 · 被引用 341 次
- Spatio-Temporal Inception Graph Convolutional Networks for Skeleton-Based Action RecognitionZhen Huang, Xu Shen, Xinmei Tian, Houqiang Li 等ACM MM 2020 · 被引用 80 次
- Towards To-a-T Spatio-Temporal Focus for Skeleton-Based Action RecognitionLipeng Ke, Kuan-Chuan Peng, Siwei LyuAAAI 2022 · 被引用 47 次
