Motion Matters: Motion-guided Modulation Network for Skeleton-based Micro-Action Recognition
Jihao Gu, Kun Li, Fei Wang, Yanyan Wei, Zhiliang Wu, Hehe Fan, Meng Wang
Abstract
Micro-Actions (MAs) are an important form of non-verbal communication in social interactions, with potential applications in human emotional analysis. However, existing methods in Micro-Action Recognition often overlook the inherent subtle changes in MAs, which limits the accuracy of distinguishing MAs with subtle changes. To address this issue, we present a novel Motion-guided Modulation Network (MMN) that implicitly captures and modulates subtle motion cues to enhance spatial-temporal representation learning. Specifically, we introduce a Motion-guided Skeletal Modulation module (MSM) to inject motion cues at the skeletal level, acting as a control signal to guide spatial representation modeling. In parallel, we design a Motion-guided Temporal Modulation module (MTM) to incorporate motion information at the frame level, facilitating the modeling of holistic motion patterns in micro-actions. Finally, we propose a motion consistency learning strategy to aggregate the motion cues from multi-scale features for micro-action classification. Experimental results on the Micro-Action 52 and iMiGUE datasets demonstrate that MMN achieves state-of-the-art performance in skeleton-based micro-action recognition, underscoring the importance of explicitly modeling subtle motion cues. The code will be available at https://github.com/momiji-bit/MMN
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- MMAD: Multi-Label Micro-Action Detection in VideosKun Li, Pengyu Liu, Dan Guo, Fei Wang et al.ICCV 2025 · 21 citations
- MA-Bench: Towards Fine-grained Micro-Action UnderstandingKun Li, Jihao Gu, Fei Wang, Zhiliang Wu et al.CVPR 2026 · 12 citations
- DLVINet: Advancing Dual-Lens Video Inpainting Beyond Parallax ConstraintsZhiliang Wu, Kun Li, Yunqiu Xu, Hehe Fan et al.AAAI 2026 · 1 citation
Builds on23
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action RecognitionYuxin Chen, Ziqi Zhang, Chunfeng Yuan, Bing Li et al.ICCV 2021 · 871 citations
Related papers
- SCD-Net: Spatiotemporal Clues Disentanglement Network for Self-Supervised Skeleton-Based Action RecognitionCong Wu, Xiao-Jun Wu, Josef Kittler, Tianyang Xu et al.AAAI 2024 · 29 citations
- Prototypical Calibrating Ambiguous Samples for Micro-Action RecognitionKun Li, Dan Guo, Guoliang Chen, Chunxiao Fan et al.AAAI 2025 · 55 citations
- Kinematic Enhanced Hypergraph Convolutional Network for Skeleton-based Human Action Recognition with LLM Training GuidesNan Ma, Beining Sun, Yiheng Han, Genbao XuACM MM 2025 · 3 citations
- Learning Multi-Granular Spatio-Temporal Graph Network for Skeleton-based Action RecognitionTailin Chen, Desen Zhou, Jian Wang, Shidong Wang et al.ACM MM 2021 · 79 citations
- Parallel Attention Interaction Network for Few-Shot Skeleton-based Action RecognitionXingyu Liu, Sanping Zhou, Le Wang, Gang HuaICCV 2023 · 17 citations
