3Mformer: Multi-order Multi-mode Transformer for Skeletal Action Recognition
Lei Wang, Piotr Koniusz
Abstract
Many skeletal action recognition models use GCNs to represent the human body by 3D body joints connected body parts. GCNs aggregate one- or few-hop graph neighbourhoods, and ignore the dependency between not linked body joints. We propose to form hypergraph to model hyperedges between graph nodes (e.g., third- and fourth-order hyper-edges capture three and four nodes) which help capture higher-order motion patterns of groups of body joints. We split action sequences into temporal blocks, Higher-order Transformer (HoT) produces embeddings of each temporal block based on (i) the body joints, (ii) pairwise links of body joints and (iii) higher-order hyper-edges of skeleton body joints. We combine such HoT embeddings of hyper-edges of orders 1,…, r by a novel Multi-order Multi-mode Transformer (3Mformer) with two modules whose order can be exchanged to achieve coupled-mode attention on coupled-mode tokens based on ‘channel-temporal block’, ‘order-channel-body joint’, ‘channel-hyper-edge (any order)’ and ‘channel-only’ pairs. The first module, called Multi-order Pooling (MP), additionally learns weighted aggregation along the hyper-edge mode, whereas the second module, Temporal block Pooling (TP), aggregates along the temporal block11For brevity, we write τ temporal blocks per sequence but τ varies. mode. Our end-to-end trainable network yields state-of-the-art results compared to GCN-, transformer- and hypergraph-based counterparts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9700f695-a84a-4e9c-aaed-a455aa44e96aCited by top-tier papers9
- Prototypical Calibrating Ambiguous Samples for Micro-Action RecognitionKun Li, Dan Guo, Guoliang Chen, Chunxiao Fan et al.AAAI 2025 · 55 citations
- LLMs are Good Action RecognizersHaoxuan Qu, Yujun Cai, Jun LiuCVPR 2024 · 37 citations
- Motion Matters: Motion-guided Modulation Network for Skeleton-based Micro-Action RecognitionJihao Gu, Kun Li, Fei Wang, Yanyan Wei et al.ACM MM 2025 · 23 citations
- Taylor Videos for Action RecognitionLei Wang, Xiuyuan Yuan, Tom Gedeon, Liang ZhengICML 2024 · 17 citations
- Frequency Guidance Matters: Skeletal Action Recognition by Frequency-Aware Mixed TransformerWenhan Wu, Ce Zheng, Zihao Yang, Chen Chen et al.ACM MM 2024 · 16 citations
Builds on25
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
- Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action RecognitionYuxin Chen, Ziqi Zhang, Chunfeng Yuan, Bing Li et al.ICCV 2021 · 871 citations
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and TextHassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang et al.NeurIPS 2021 · 782 citations
Related papers
- Hierarchically Decomposed Graph Convolutional Networks for Skeleton-Based Action RecognitionJungho Lee, Minhyeok Lee, Dogyoon Lee, Sangyoun LeeICCV 2023 · 236 citations
- Spatio-Temporal Fusion for Human Action Recognition via Joint Trajectory GraphYaolin Zheng, Hongbo Huang, Xiuying Wang, Xiaoxu Yan et al.AAAI 2024 · 22 citations
- Skeleton MixFormer: Multivariate Topology Representation for Skeleton-based Action RecognitionWentian Xin, Qiguang Miao, Yi Liu, Ruyi Liu et al.ACM MM 2023 · 66 citations
- Part-Level Graph Convolutional Network for Skeleton-Based Action RecognitionLinjiang Huang, Yan Huang, Wanli Ouyang, Liang WangAAAI 2020 · 111 citations
- HyperGait: Unleashing the Power of Parsing for Gait Recognition in the Wild via HypergraphJinkai Zheng, Jiaqing Wei, Xinxiang Jin, Yaoqi Sun et al.CVPR 2026
