Active Vision for Early Recognition of Human Actions
Boyu Wang, Lihan Huang, Minh Hoai
Abstract
We propose a method for early recognition of human actions, one that can take advantages of multiple cameras while satisfying the constraints due to limited communication bandwidth and processing power. Our method considers multiple cameras, and at each time step, it will decide the best camera to use so that a confident recognition decision can be reached as soon as possible. We formulate the camera selection problem as a sequential decision process, and learn a view selection policy based on reinforcement learning. We also develop a novel recurrent neural network architecture to account for the unobserved video frames and the irregular intervals between the observed frames. Experiments on three datasets demonstrate the effectiveness of our approach for early recognition of human actions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- RAIN: Reinforced Hybrid Attention Inference Network for Motion ForecastingJiachen Li, Fan Yang, Hengbo Ma, Srikanth Malla et al.ICCV 2021 · 49 citations
- The Power of Log-Sum-Exp: Sequential Density Ratio Matrix Estimation for Speed-Accuracy OptimizationTaiki Miyagawa, Akinori F. EbiharaICML 2021 · 4 citations
- MoBind: Motion Binding for Fine-Grained IMU-Video Pose AlignmentDuc Duy Nguyen, Tat-Jun Chin, Minh HoaiCVPR 2026 · 1 citation
Builds on2
Related papers
- Ego-Pose Estimation and Forecasting As Real-Time PD ControlYe Yuan, Kris KitaniICCV 2019 · 147 citations
- Learning to Select Views for Efficient Multi-View UnderstandingYunzhong Hou, Stephen Gould, Liang ZhengCVPR 2024
- Temporal Recurrent Networks for Online Action DetectionMingze Xu, Mingfei Gao, Yi-Ting Chen, Larry Davis et al.ICCV 2019 · 201 citations
- Learning Active Camera for Multi-Object NavigationPeihao Chen, Dongyu Ji, Kunyang Lin, Weiwen Hu et al.NeurIPS 2022 · 40 citations
- Multi-Agent Reinforcement Learning Based Frame Sampling for Effective Untrimmed Video RecognitionWenhao Wu, Dongliang He, Xiao Tan, Shifeng Chen et al.ICCV 2019 · 135 citations
