Otter: Mitigating Background Distractions of Wide-Angle Few-Shot Action Recognition with Enhanced RWKV
Wenbo Huang, Jinghui Zhang, Zhenghao Chen, Guang Li, Lei Zhang, Yang Cao, Fang Dong, Takahiro Ogawa, Miki Haseyama
Abstract
Wide-angle videos in few-shot action recognition (FSAR) effectively express actions within specific scenarios. However, without a global understanding of both subjects and background, recognizing actions in such samples remains challenging because of the background distractions. Receptance Weighted Key Value (RWKV), which learns interaction between various dimensions, shows promise for global modeling. While directly applying RWKV to wide-angle FSAR may fail to highlight subjects due to excessive background information. Additionally, temporal relation degraded by frames with similar backgrounds is difficult to reconstruct, further impacting performance. Therefore, we design the CompOund SegmenTation and Temporal REconstructing RWKV (Otter). Specifically, the Compound Segmentation Module (CSM) is devised to segment and emphasize key patches in each frame, effectively highlighting subjects against background information. The Temporal Reconstruction Module (TRM) is incorporated into the temporal-enhanced prototype construction to enable bidirectional scanning, allowing better reconstruct temporal relation. Furthermore, a regular prototype is combined with the temporal-enhanced prototype to simultaneously enhance subject emphasis and temporal modeling, improving wide-angle FSAR performance. Extensive experiments on benchmarks such as SSv2, Kinetics, UCF101, and HMDB51 demonstrate that Otter achieves state-of-the-art performance. Extra evaluation on the VideoBadminton dataset further validates the superiority of Otter in wide-angle FSAR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 66209764-7da6-4f1b-aebf-3ee2b1a2e40fBuilds on19
- Collect and Select: Semantic Alignment Metric Learning for Few-Shot LearningFusheng Hao, Fengxiang He, Jun Cheng, Lei Wang et al.ICCV 2019 · 146 citations
- TA2N: Two-Stage Action Alignment Network for Few-Shot Action RecognitionShuyuan Li, Huabin Liu, Rui Qian, Yuxi Li et al.AAAI 2022 · 98 citations
- Depth Guided Adaptive Meta-Fusion Network for Few-shot Video RecognitionYuqian Fu, Li Zhang, Junke Wang, Yanwei Fu et al.ACM MM 2020 · 97 citations
- Feature Prediction Diffusion Model for Video Anomaly DetectionCheng Yan, Shiyu Zhang, Yang Liu, Guansong Pang et al.ICCV 2023 · 76 citations
- Motion-modulated Temporal Fragment Alignment Network For Few-Shot Action RecognitionJiamin Wu, Tianzhu Zhang, Zhe Zhang, Feng Wu et al.CVPR 2022 · 73 citations
Related papers
- Spatio-temporal Relation Modeling for Few-shot Action RecognitionAnirudh Thatipelli, Sanath Narayan, Salman Khan, Rao Muhammad Anwer et al.CVPR 2022 · 144 citations
- Temporal-Relational CrossTransformers for Few-Shot Action RecognitionToby Perrett, Alessandro Masullo, Tilo Burghardt, Majid Mirmehdi et al.CVPR 2021
- SOAP: Enhancing Spatio-Temporal Relation and Motion Information Capturing for Few-Shot Action RecognitionWenbo Huang, Jinghui Zhang, Xuwei Qian, Zhen Wu et al.ACM MM 2024 · 8 citations
- Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action RecognitionPulkit Kumar, Shuaiyi Huang, Matthew Walmer, Sai Saketh Rambhatla et al.ICCV 2025
- Revisiting the Spatial and Temporal Modeling for Few-Shot Action RecognitionJiazheng Xing, Mengmeng Wang, Yong Liu, Boyu MuAAAI 2023 · 51 citations
