Counting Out Time: Class Agnostic Video Repetition Counting in the Wild
Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson, Pierre Sermanet, Andrew Zisserman
Abstract
We present an approach for estimating the period with which an action is repeated in a video. The crux of the approach lies in constraining the period prediction module to use temporal self-similarity as an intermediate representation bottleneck that allows generalization to unseen repetitions in videos in the wild. We train this model, called RepNet, with a synthetic dataset that is generated from a large unlabeled video collection by sampling short clips of varying lengths and repeating them with different periods and counts. This combination of synthetic data and a powerful yet constrained model, allows us to predict periods in a class-agnostic fashion. Our model substantially exceeds the state of the art performance on existing periodicity (PERTUBE) and repetition counting (QUVA) benchmarks. We also collect a new challenging dataset called Countix (∼90 times larger than existing datasets) which captures the challenges of repetition counting in real-world videos. Project webpage: https://sites.google. com/view/repnet .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 36f5b5ba-1772-419c-afbf-561ea8b85e29Cited by top-tier papers24
- The Way to my Heart is through Contrastive Learning: Remote Photoplethysmography from Unlabelled VideoJohn Gideon, Simon StentICCV 2021 · 153 citations
- Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and GroundingChristopher Clark, Jieyu Zhang, Zixian Ma, Jae Sung Park et al.CVPR 2026 · 144 citations
- Knowing Where to Focus: Event-aware Transformer for Video GroundingJinhyun Jang, Jungin Park, Jin Kim, Hyeongjun Kwon et al.ICCV 2023 · 103 citations
- Zero-shot Natural Language Video LocalizationJinwoo Nam, Daechul Ahn, Dongyeop Kang, Seong Jong Ha et al.ICCV 2021 · 60 citations
- TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action CountingHuazhang Hu, Sixun Dong, Yiqun Zhao, Dongze Lian et al.CVPR 2022 · 57 citations
Builds on1
Related papers
- Context-Aware and Scale-Insensitive Temporal Repetition CountingHuaidong Zhang, Xuemiao Xu, Guoqiang Han, Shengfeng HeCVPR 2020
- Learning Self-Similarity in Space and Time as Generalized Motion for Video Action RecognitionHeeseung Kwon, Manjin Kim, Suha Kwak, Minsu ChoICCV 2021 · 49 citations
- CountLLM: Towards Generalizable Repetitive Action Counting via Large Language ModelZiyu Yao, Xuxin Cheng, Zhiqi Huang, Lei LiCVPR 2025
- Open-World Object Counting in VideosNiki Amini-Naieni, Andrew ZissermanAAAI 2026 · 6 citations
- Can't make an Omelette without Breaking some Eggs: Plausible Action Anticipation using Large Video-Language ModelsHimangi Mittal, Nakul Agarwal, Shao-Yuan Lo, Kwonjoon LeeCVPR 2024 · 14 citations
