TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action Counting
Huazhang Hu, Sixun Dong, Yiqun Zhao, Dongze Lian, Zhengxin Li, Shenghua Gao
Abstract
Counting repetitive actions are widely seen in human activities such as physical exercise. Existing methods focus on performing repetitive action counting in short videos, which is tough for dealing with longer videos in more realistic scenarios. In the data-driven era, the degradation of such generalization capability is mainly attributed to the lack of long video datasets. To complement this margin, we introduce a new large-scale repetitive action counting dataset covering a wide variety of video lengths, along with more realistic situations where action interruption or action inconsistencies occur in the video. Besides, we also provide a fine-grained annotation of the action cycles instead of just counting annotation along with a numerical value. Such a dataset contains 1,451 videos with about 20,000 annotations, which is more challenging. For repetitive action counting towards more realistic scenarios, we further propose encoding multi-scale temporal correlation with transformers that can take into account both performance and efficiency. Furthermore, with the help of fine-grained annotation of action cycles, we propose a density map regression-based method to predict the action period, which yields better performance with sufficient interpretability. Our proposed method outperforms state-of-the-art methods on all datasets and also achieves better performance on the unseen dataset without fine-tuning. The dataset and code are available <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> https://svip-lab.github.io/dataset/RepCount_dataset.html.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- RhythmMamba: Fast, Lightweight, and Accurate Remote Physiological MeasurementBochao Zou, Zizheng Guo, Xiaocheng Hu, Huimin MaAAAI 2025 · 24 citations
- AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMsLidong Lu, Guo Chen, Zhu Wei, Zhiqi Li et al.CVPR 2026 · 23 citations
- Tubelet-Contrastive Self-Supervision for Video-Efficient GeneralizationFida Mohammad Thoker, Hazel Doughty, Cees G. M. SnoekICCV 2023 · 13 citations
- Count What You Want: Exemplar Identification and Few-Shot Counting of Human Actions in the WildYifeng Huang, Duc Duy Nguyen, Lam Nguyen, Cuong Pham et al.AAAI 2024 · 5 citations
- Mavors: Multi-granularity Video Representation for Multimodal Large Language ModelYang Shi, Jiaheng Liu, Yushuo Guan, Zhenhua Wu et al.ACM MM 2025 · 1 citation
Builds on9
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- Adaptive Density Map Generation for Crowd CountingJia Wan, Antoni B. ChanICCV 2019 · 171 citations
Related papers
- Context-Aware and Scale-Insensitive Temporal Repetition CountingHuaidong Zhang, Xuemiao Xu, Guoqiang Han, Shengfeng HeCVPR 2020
- Counting Out Time: Class Agnostic Video Repetition Counting in the WildDebidatta Dwibedi, Yusuf Aytar, Jonathan Tompson, Pierre Sermanet et al.CVPR 2020
- Repetitive Activity Counting by Sight and SoundYunhua Zhang, Ling Shao, Cees G. M. SnoekCVPR 2021
- CountLLM: Towards Generalizable Repetitive Action Counting via Large Language ModelZiyu Yao, Xuxin Cheng, Zhiqi Huang, Lei LiCVPR 2025
- MS-TCT: Multi-Scale Temporal ConvTransformer for Action DetectionRui Dai, Srijan Das, Kumara Kahatapitiya, Michael S. Ryoo et al.CVPR 2022 · 93 citations
