RTMO: Towards High-Performance One-Stage Real-Time Multi-Person Pose Estimation
Peng Lu, Tao Jiang, Yining Li, Xiangtai Li, Kai Chen, Wenming Yang
摘要
Real-time multi-person pose estimation presents signif-icant challenges in balancing speed and precision. While two-stage top-down methods slow down as the number of people in the image increases, existing one-stage meth-ods often fail to simultaneously deliver high accuracy and real-time performance. This paper introduces RTMO, a one-stage pose estimation framework that seamlessly inte-grates coordinate classification by representing keypoints using dual I-D heatmaps within the YOLO architecture, achieving accuracy comparable to top-down methods while maintaining high speed. We propose a dynamic coordi-nate classifier and a tailored loss function for heatmap learning, specifically designed to address the incompati-bilities between coordinate classification and dense pre-diction models. RTMO outperforms state-of-the-art one-stage pose estimators, achieving 1.1% higher AP on COCO while operating about 9 times faster with the same back-bone. Our largest model, RTMO-1, attains 74.8% AP on COCO va12017 and 141 FPS on a single V100 GPU, demonstrating its efficiency and accuracy. The code and models are available at https://github.com/open-mmlab/mmpose/tree/main/projects/rtmo.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Referring Human Pose and Mask Estimation In the WildBo Miao, Mingtao Feng, Zijie Wu, Mohammed Bennamoun 等NeurIPS 2024 · 被引用 12 次
- EDM: Efficient Deep Feature MatchingXi Li, Tong Rao, Cihui PanICCV 2025 · 被引用 6 次
- FIP: Endowing Robust Motion Capture on Daily Garment by Fusing Flex and Inertial SensorsRuonan Zheng, Jiawei Fang, Yuan Yao, Xiaoxia Gao 等CHI 2025 · 被引用 5 次
- MultiCam: On-the-fly Multi-Camera Pose Estimation Using Spatiotemporal Overlaps of Known ObjectsShiyu Li, Hannah Schieber, Kristoffer Waldow, Benjamin Busam 等IEEE VR 2026 · 被引用 2 次
- End-to-End Multi-Person Pose Estimation with Pose-Aware Video TransformerYonghui Yu, Jiahang Cai, Xun Wang, Wenwu YangAAAI 2026 · 被引用 2 次
它引用的顶会 Paper19
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi 等ICCV 2019 · 被引用 3,348 次
- DETRs Beat YOLOs on Real-time Object DetectionYian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei 等CVPR 2024 · 被引用 3,046 次
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 被引用 1,105 次
相关 Paper
- InsPose: Instance-Aware Networks for Single-Stage Multi-Person Pose EstimationDahu Shi, Xing Wei, Xiaodong Yu, Wenming Tan 等ACM MM 2021 · 被引用 40 次
- FCPose: Fully Convolutional Multi-Person Pose Estimation With Dynamic Instance-Aware ConvolutionsWeian Mao, Zhi Tian, Xinlong Wang, Chunhua ShenCVPR 2021
- Single-Network Whole-Body Pose EstimationGines Hidalgo Martinez, Yaadhav Raaj, Haroon Idrees, Donglai Xiang 等ICCV 2019 · 被引用 115 次
- SIMPLE: SIngle-network with Mimicking and Point Learning for Bottom-up Human Pose EstimationJiabin Zhang, Zheng Zhu, Jiwen Lu, Junjie Huang 等AAAI 2021 · 被引用 14 次
- Learning Local-Global Contextual Adaptation for Multi-Person Pose EstimationNan Xue, Tianfu Wu, Gui-Song Xia, Liangpei ZhangCVPR 2022 · 被引用 42 次
