Count What You Want: Exemplar Identification and Few-Shot Counting of Human Actions in the Wild
Yifeng Huang, Duc Duy Nguyen, Lam Nguyen, Cuong Pham, Minh Hoai
摘要
This paper addresses the task of counting human actions of interest using sensor data from wearable devices. We propose a novel exemplar-based framework, allowing users to provide exemplars of the actions they want to count by vocalizing predefined sounds one'', two'', and ``three''. Our method first localizes temporal positions of these utterances from the audio sequence. These positions serve as the basis for identifying exemplars representing the action class of interest. A similarity map is then computed between the exemplars and the entire sensor data sequence, which is further fed into a density estimation module to generate a sequence of estimated density values. Summing these density values provides the final count. To develop and evaluate our approach, we introduce a diverse and realistic dataset consisting of real-world data from 37 subjects and 50 action categories, encompassing both sensor and audio data. The experiments on this dataset demonstrate the viability of the proposed method in counting instances of actions from new classes and subjects that were not part of the training data. On average, the discrepancy between the predicted count and the ground truth value is 7.47, significantly lower than the errors of the frequency-based and transformer-based methods. Our project, code and dataset can be found at https://github.com/cvlab-stonybrook/ExRAC.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- MoBind: Motion Binding for Fine-Grained IMU-Video Pose AlignmentDuc Duy Nguyen, Tat-Jun Chin, Minh HoaiCVPR 2026 · 被引用 1 次
- Decoupling What to Count and Where to See for Referring Expression CountingYuda Zou, Zijian Zhang, Yongchao XuAAAI 2026
它引用的顶会 Paper13
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang 等AAAI 2021 · 被引用 7,289 次
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 被引用 5,824 次
- Pyraformer: Low-Complexity Pyramidal Attention for Long-Range Time Series Modeling and ForecastingShizhan Liu, Hang Yu, Cong Liao, Jianguo Li 等ICLR 2022 · 被引用 975 次
- Crowd Counting with Decomposed UncertaintyMin-hwan Oh, Peder A. Olsen, Karthikeyan Natesan RamamurthyAAAI 2020 · 被引用 118 次
- Represent, Compare, and Learn: A Similarity-Aware Framework for Class-Agnostic CountingMin Shi, Hao Lu, Chen Feng, Chengxin Liu 等CVPR 2022 · 被引用 99 次
相关 Paper
- TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action CountingHuazhang Hu, Sixun Dong, Yiqun Zhao, Dongze Lian 等CVPR 2022 · 被引用 57 次
- Learning To Count EverythingViresh Ranjan, Udbhav Sharma, Thu Nguyen, Minh HoaiCVPR 2021
- MoCount: Motion-Based Repetitive Action CountingRuocheng Gu, Sen Jia, Yule Ma, Jinqin Zhong 等ACM MM 2025 · 被引用 4 次
- Repetitive Activity Counting by Sight and SoundYunhua Zhang, Ling Shao, Cees G. M. SnoekCVPR 2021
- Action Dubber: Timing Audible Actions via Inflectional FlowWenlong Wan, Weiying Zheng, Tianyi Xiang, Guiqing Li 等ICML 2025
