Deep Reinforcement Learning for Active Human Pose Estimation
Erik Gärtner, Aleksis Pirinen, Cristian Sminchisescu
Abstract
Most 3d human pose estimation methods assume that input -be it images of a scene collected from one or several viewpoints, or from a video -is given. Consequently, they focus on estimates leveraging prior knowledge and measurement by fusing information spatially and/or temporally, whenever available. In this paper we address the problem of an active observer with freedom to move and explore the scene spatially -in 'time-freeze' mode -and/or temporally, by selecting informative viewpoints that improve its estimation accuracy. Towards this end, we introduce Pose-DRL, a fully trainable deep reinforcement learning-based active pose estimation architecture which learns to select appropriate views, in space and time, to feed an underlying monocular pose estimator. We evaluate our model using single-and multi-target estimators with strong result in both settings. Our system further learns automatic stopping conditions in time and transition functions to the next temporal processing step in videos. In extensive experiments with the Panoptic multi-view setup, and for complex scenes containing multiple people, we show that our model learns to select viewpoints that yield significantly more accurate pose estimates compared to strong multi-view baselines. Code is available: https://github.com/aleksispi/pose-drl . * Denotes equal contribution, order determined by coin flip.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2604f290-e5cc-4cbf-b0f0-1f50612c98daCited by top-tier papers6
- Meta Agent Teaming Active Learning for Pose EstimationJia Gong, Zhipeng Fan, Qiuhong Ke, Hossein Rahmani et al.CVPR 2022 · 53 citations
- Embodied Visual Active Learning for Semantic SegmentationDavid Nilsson, Aleksis Pirinen, Erik Gärtner, Cristian SminchisescuAAAI 2021 · 37 citations
- H3WB: Human3.6M 3D WholeBody Dataset and BenchmarkYue Zhu, Nermin Samet, David PicardICCV 2023 · 34 citations
- Efficient Virtual View Selection for 3D Hand Pose EstimationJian Cheng, Yanguang Wan, Dexin Zuo, Cuixia Ma et al.AAAI 2022 · 30 citations
- Glimpse-Attend-and-Explore: Self-Attention for Active Visual ExplorationSoroush Seifi, Abhishek Jha, Tinne TuytelaarsICCV 2021 · 11 citations
Related papers
- ActiveMoCap: Optimized Viewpoint Selection for Active Human Motion CaptureSena Kiciroglu, Helge Rhodin, Sudipta N. Sinha, Mathieu Salzmann et al.CVPR 2020
- PFRL: Pose-Free Reinforcement Learning for 6D Pose EstimationJianzhun Shao, Yuhang Jiang, Gu Wang, Zhigang Li et al.CVPR 2020
- Ego-Pose Estimation and Forecasting As Real-Time PD ControlYe Yuan, Kris KitaniICCV 2019 · 147 citations
- TEMPO: Efficient Multi-View Pose Estimation, Tracking, and ForecastingRohan Choudhury, Kris M. Kitani, László A. JeniICCV 2023 · 32 citations
- DSP: Dense-Sparse Parallel Networks for Self-supervised 3D Multi-person Pose Estimation from Multiple ViewsYang Liu, Zhiyong ZhangACM MM 2025
