Multi-Agent Reinforcement Learning Based Frame Sampling for Effective Untrimmed Video Recognition
Wenhao Wu, Dongliang He, Xiao Tan, Shifeng Chen, Shilei Wen
Abstract
Video Recognition has drawn great research interest and great progress has been made. A suitable frame sampling strategy can improve the accuracy and efficiency of recognition. However, mainstream solutions generally adopt hand-crafted frame sampling strategies for recognition. It could degrade the performance, especially in untrimmed videos, due to the variation of frame-level saliency. To this end, we concentrate on improving untrimmed video classification via developing a learning-based frame sampling strategy. We intuitively formulate the frame sampling procedure as multiple parallel Markov decision processes, each of which aims at picking out a frame/clip by gradually adjusting an initial sampling. Then we propose to solve the problems with multi-agent reinforcement learning (MARL). Our MARL framework is composed of a novel RNN-based context-aware observation network which jointly models context information among nearby agents and historical states of a specific agent, a policy network which generates the probability distribution over a predefined action space at each step and a classification network for reward calculation as well as final recognition. Extensive experimental results show that our MARL-based scheme remarkably outperforms hand-crafted strategies with various 2D and 3D baseline methods. Our single RGB model achieves a comparable performance of ActivityNet v1.3 champion submission with multi-modal multi-model fusion and new state-of-the-art results on YouTube Birds and YouTube Cars.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2413f45f-559c-4330-8762-9d4f9275e30cCited by top-tier papers45
- Self-Chained Image-Language Model for Video Localization and Question AnsweringShoubin Yu, Jaemin Cho, Prateek Yadav, Mohit BansalNeurIPS 2023 · 281 citations
- Location-Aware Graph Convolutional Networks for Video Question AnsweringDeng Huang, Peihao Chen, Runhao Zeng, Qing Du et al.AAAI 2020 · 187 citations
- SMART Frame Selection for Action RecognitionShreyank N. Gowda, Marcus Rohrbach, Laura Sevilla-LaraAAAI 2021 · 171 citations
- Revisiting Classifier: Transferring Vision-Language Models for Video RecognitionWenhao Wu, Zhun Sun, Wanli OuyangAAAI 2023 · 141 citations
- Tree-Structured Policy Based Progressive Reinforcement Learning for Temporally Language Grounding in VideoJie Wu, Guanbin Li, Si Liu, Liang LinAAAI 2020 · 117 citations
Related papers
- An Efficient Framework for Dense Video CaptioningMaitreya Suin, A. N. RajagopalanAAAI 2020 · 49 citations
- Multi-Instance Multi-Label Action Recognition and Localization Based on Spatio-Temporal Pre-Trimming for Untrimmed VideosXiaoyu Zhang, Haichao Shi, Changsheng Li, Peng LiAAAI 2020 · 37 citations
- Annotation-Efficient Untrimmed Video Action RecognitionYixiong Zou, Shanghang Zhang, Guangyao Chen, Yonghong Tian et al.ACM MM 2021 · 7 citations
- AdaMML: Adaptive Multi-Modal Learning for Efficient Video RecognitionRameswar Panda, Chun-Fu (Richard) Chen, Quanfu Fan, Ximeng Sun et al.ICCV 2021 · 65 citations
- Straight to the Point: Fast-Forwarding Videos via Reinforcement Learning Using Textual DataWashington L. S. Ramos, Michel Melo Silva, Edson R. Araujo, Leandro Soriano Marcolino et al.CVPR 2020
