Repetitive Activity Counting by Sight and Sound
Yunhua Zhang, Ling Shao, Cees G. M. Snoek
Abstract
This paper strives for repetitive activity counting in videos. Different from existing works, which all analyze the visual video content only, we incorporate for the first time the corresponding sound into the repetition counting process. This benefits accuracy in challenging vision conditions such as occlusion, dramatic camera view changes, low resolution, etc. We propose a model that starts with analyzing the sight and sound streams separately. Then an audiovisual temporal stride decision module and a reliability estimation module are introduced to exploit cross-modal temporal interaction. For learning and evaluation, an existing dataset is repurposed and reorganized to allow for repetition counting with sight and sound. We also introduce a variant of this dataset for repetition counting under challenging vision conditions. Experiments demonstrate the benefit of sound, as well as the other introduced modules, for repetition counting. Our sight-only model already outperforms the state-of-the-art by itself, when we add sound, results improve notably, especially under harsh vision conditions. The code and datasets are available at https://github.com/xiaobai1217/ RepetitionCounting .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4cb28673-ae53-4bb7-9d16-dee4a470e155Cited by top-tier papers15
- TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action CountingHuazhang Hu, Sixun Dong, Yiqun Zhao, Dongze Lian et al.CVPR 2022 · 57 citations
- Read, Watch and Scream! Sound Generation from Text and VideoYujin Jeong, Yunji Kim, Sanghyuk Chun, Jiyoung LeeAAAI 2025 · 48 citations
- Audio-Adaptive Activity Recognition Across Video DomainsYunhua Zhang, Hazel Doughty, Ling Shao, Cees G. M. SnoekCVPR 2022 · 31 citations
- AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMsLidong Lu, Guo Chen, Zhu Wei, Zhiqi Li et al.CVPR 2026 · 23 citations
- Learning Unseen Modality InteractionYunhua Zhang, Hazel Doughty, Cees SnoekNeurIPS 2023 · 16 citations
Builds on11
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 271 citations
- SCSampler: Sampling Salient Clips From Video for Efficient Action RecognitionBruno Korbar, Du Tran, Lorenzo TorresaniICCV 2019 · 257 citations
- Dual Attention Matching for Audio-Visual Event LocalizationYu Wu, Linchao Zhu, Yan Yan, Yi YangICCV 2019 · 233 citations
- Co-Separating Sounds of Visual ObjectsRuohan Gao, Kristen GraumanICCV 2019 · 224 citations
- Recursive Visual Sound Separation Using Minus-Plus NetXudong Xu, Bo Dai, Dahua LinICCV 2019 · 95 citations
Related papers
- Context-Aware and Scale-Insensitive Temporal Repetition CountingHuaidong Zhang, Xuemiao Xu, Guoqiang Han, Shengfeng HeCVPR 2020
- Exploring Heterogeneous Clues for Weakly-Supervised Audio-Visual Video ParsingYu Wu, Yi YangCVPR 2021
- CACE-Net: Co-guidance Attention and Contrastive Enhancement for Effective Audio-Visual Event LocalizationXiang He, Xiangxi Liu, Yang Li, Dongcheng Zhao et al.ACM MM 2024 · 8 citations
- Dense-Localizing Audio-Visual Events in Untrimmed Videos: A Large-Scale Benchmark and BaselineTiantian Geng, Teng Wang, Jinming Duan, Runmin Cong et al.CVPR 2023
- Self-Supervised Moving Vehicle Tracking With Stereo SoundChuang Gan, Hang Zhao, Peihao Chen, David D. Cox et al.ICCV 2019 · 157 citations
