REMOTE: Reinforced Motion Transformation Network for Semi-supervised 2D Pose Estimation in Videos
Xianzheng Ma, Hossein Rahmani, Zhipeng Fan, Bin Yang, Jun Chen, Jun Liu
摘要
Existing approaches for 2D pose estimation in videos often require a large number of dense annotations, which are costly and labor intensive to acquire. In this paper, we propose a semi-supervised REinforced MOtion Transformation nEtwork (REMOTE) to leverage a few labeled frames and temporal pose variations in videos, which enables effective learning of 2D pose estimation in sparsely annotated videos. Specifically, we introduce a Motion Transformer (MT) module to perform cross frame reconstruction, aiming to learn motion dynamic knowledge in videos. Besides, a novel reinforcement learning-based Frame Selection Agent (FSA) is designed within our framework, which is able to harness informative frame pairs on the fly to enhance the pose estimator under our cross reconstruction mechanism. We conduct extensive experiments that show the efficacy of our proposed REMOTE framework.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in VideoJinlu Zhang, Zhigang Tu, Jianyu Yang, Yujin Chen 等CVPR 2022 · 被引用 356 次
- CUPS: Improving Human Pose-Shape Estimators with Conformalized Deep UncertaintyHarry Zhang, Luca CarloneICML 2025
- CHAMP: Conformalized 3D Human Multi-Hypothesis Pose EstimatorsHarry Zhang, Luca CarloneICLR 2025
它引用的顶会 Paper9
- Human Motion Prediction via Spatio-Temporal InpaintingAlejandro Hernandez Ruiz, Jürgen Gall, Francesc MorenoICCV 2019 · 被引用 233 次
- Multi-Agent Reinforcement Learning Based Frame Sampling for Effective Untrimmed Video RecognitionWenhao Wu, Dongliang He, Xiao Tan, Shifeng Chen 等ICCV 2019 · 被引用 135 次
- Reinforced active learning for image segmentationArantxa Casanova, Pedro O. Pinheiro, Negar Rostamzadeh, Christopher J. PalICLR 2020 · 被引用 127 次
- Adaptive Focus for Efficient Video RecognitionYulin Wang, Zhaoxi Chen, Haojun Jiang, Shiji Song 等ICCV 2021 · 被引用 117 次
- Dynamic Kernel Distillation for Efficient Pose Estimation in VideosXuecheng Nie, Yuncheng Li, Linjie Luo, Ning Zhang 等ICCV 2019 · 被引用 76 次
相关 Paper
- SpatioTemporal Learning for Human Pose Estimation in Sparsely-Labeled VideosYingying Jiao, Zhigang Wang, Sifan Wu, Shaojing Fan 等AAAI 2025 · 被引用 5 次
- Joint Inductive and Transductive Learning for Video Object SegmentationYunyao Mao, Ning Wang, Wengang Zhou, Houqiang LiICCV 2021 · 被引用 111 次
- Siamese Network with Interactive Transformer for Video Object SegmentationMeng Lan, Jing Zhang, Fengxiang He, Lefei ZhangAAAI 2022 · 被引用 41 次
- PFRL: Pose-Free Reinforcement Learning for 6D Pose EstimationJianzhun Shao, Yuhang Jiang, Gu Wang, Zhigang Li 等CVPR 2020
- SSTVOS: Sparse Spatiotemporal Transformers for Video Object SegmentationBrendan Duke, Abdalla Ahmed, Christian Wolf, Parham Aarabi 等CVPR 2021
