MoCount: Motion-Based Repetitive Action Counting
Ruocheng Gu, Sen Jia, Yule Ma, Jinqin Zhong, Jenq-Neng Hwang, Lei Li
Abstract
Existing action counting methods typically rely on pixel-based changes within videos, leading to high computational redundancy and low accuracy due to the limited spatial sensitivity. To address these challenges, we introduce MoCount, the first framework that leverages 3D motion representations for counting tasks. MoCount significantly reduces computational overhead and improves counting accuracy, benefiting from the simplicity of motion representation and strong spatial sensitivity. Specifically, we utilize a motion estimator to convert video subjects into 3D motion data. A motion encoder, combined with a Sparse Spatial-Temporal module, is then applied to extract robust human body representations, yielding precise counting results. Extensive experiments on the RepCount and UCFRep datasets show that MoCount achieves state-of-the-art performance, reducing inference latency by approximately 2-3 times compared to existing video counting models. These advantages position MoCount as a leading solution for real-world action counting applications.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b2a308cd-118c-4f18-8be1-da6032dd3d6aCited by top-tier papers11
- GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal GenerationTianchen Deng, Xuefeng Chen, Yi Chen, Qu Chen et al.CVPR 2026 · 31 citations
- ReTrack: Evidence-Driven Dual-Stream Directional Anchor Calibration Network for Composed Video RetrievalZixu Li, Yupeng Hu, Zhiwei Chen, Qinlei Huang et al.AAAI 2026 · 24 citations
- Self-Destructive Language ModelsYuhui Wang, Rongyi Zhu, Ting WangICLR 2026 · 14 citations
- Reasoning or Retrieval? A Study of Answer Attribution on Large Reasoning ModelsYuhui Wang, Changjiang Li, Guangke Chen, Jiacheng Liang et al.ICLR 2026 · 13 citations
- RAM: Recover Any 3D Human Motion in-the-WildSen Jia, Ning Zhu, Jinqin Zhong, Jiale Zhou et al.CVPR 2026 · 12 citations
Related papers
- Context-Aware and Scale-Insensitive Temporal Repetition CountingHuaidong Zhang, Xuemiao Xu, Guoqiang Han, Shengfeng HeCVPR 2020
- Deep Analysis of CNN-Based Spatio-Temporal Representations for Action RecognitionChun-Fu (Richard) Chen, Rameswar Panda, Kandan Ramakrishnan, Rogério Feris et al.CVPR 2021
- Unified Keypoint-Based Action Recognition Framework via Structured Keypoint PoolingRyo Hachiuma, Fumiaki Sato, Taiki SekiiCVPR 2023
- ACTION-Net: Multipath Excitation for Action RecognitionZhengwei Wang, Qi She, Aljosa SmolicCVPR 2021
- 3D Human Pose Estimation Using Spatio-Temporal Networks with Explicit Occlusion TrainingYu Cheng, Bo Yang, Bo Wang, Robby T. TanAAAI 2020 · 145 citations
