MoCount: Motion-Based Repetitive Action Counting
Ruocheng Gu, Sen Jia, Yule Ma, Jinqin Zhong, Jenq-Neng Hwang, Lei Li
摘要
Existing action counting methods typically rely on pixel-based changes within videos, leading to high computational redundancy and low accuracy due to the limited spatial sensitivity. To address these challenges, we introduce MoCount, the first framework that leverages 3D motion representations for counting tasks. MoCount significantly reduces computational overhead and improves counting accuracy, benefiting from the simplicity of motion representation and strong spatial sensitivity. Specifically, we utilize a motion estimator to convert video subjects into 3D motion data. A motion encoder, combined with a Sparse Spatial-Temporal module, is then applied to extract robust human body representations, yielding precise counting results. Extensive experiments on the RepCount and UCFRep datasets show that MoCount achieves state-of-the-art performance, reducing inference latency by approximately 2-3 times compared to existing video counting models. These advantages position MoCount as a leading solution for real-world action counting applications.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper11
- GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal GenerationTianchen Deng, Xuefeng Chen, Yi Chen, Qu Chen 等CVPR 2026 · 被引用 31 次
- ReTrack: Evidence-Driven Dual-Stream Directional Anchor Calibration Network for Composed Video RetrievalZixu Li, Yupeng Hu, Zhiwei Chen, Qinlei Huang 等AAAI 2026 · 被引用 24 次
- Self-Destructive Language ModelsYuhui Wang, Rongyi Zhu, Ting WangICLR 2026 · 被引用 14 次
- Reasoning or Retrieval? A Study of Answer Attribution on Large Reasoning ModelsYuhui Wang, Changjiang Li, Guangke Chen, Jiacheng Liang 等ICLR 2026 · 被引用 13 次
- RAM: Recover Any 3D Human Motion in-the-WildSen Jia, Ning Zhu, Jinqin Zhong, Jiale Zhou 等CVPR 2026 · 被引用 12 次
相关 Paper
- Context-Aware and Scale-Insensitive Temporal Repetition CountingHuaidong Zhang, Xuemiao Xu, Guoqiang Han, Shengfeng HeCVPR 2020
- Deep Analysis of CNN-Based Spatio-Temporal Representations for Action RecognitionChun-Fu (Richard) Chen, Rameswar Panda, Kandan Ramakrishnan, Rogério Feris 等CVPR 2021
- Unified Keypoint-Based Action Recognition Framework via Structured Keypoint PoolingRyo Hachiuma, Fumiaki Sato, Taiki SekiiCVPR 2023
- ACTION-Net: Multipath Excitation for Action RecognitionZhengwei Wang, Qi She, Aljosa SmolicCVPR 2021
- 3D Human Pose Estimation Using Spatio-Temporal Networks with Explicit Occlusion TrainingYu Cheng, Bo Yang, Bo Wang, Robby T. TanAAAI 2020 · 被引用 145 次
