RTMO: Towards High-Performance One-Stage Real-Time Multi-Person Pose Estimation
Peng Lu, Tao Jiang, Yining Li, Xiangtai Li, Kai Chen, Wenming Yang
Abstract
Real-time multi-person pose estimation presents signif-icant challenges in balancing speed and precision. While two-stage top-down methods slow down as the number of people in the image increases, existing one-stage meth-ods often fail to simultaneously deliver high accuracy and real-time performance. This paper introduces RTMO, a one-stage pose estimation framework that seamlessly inte-grates coordinate classification by representing keypoints using dual I-D heatmaps within the YOLO architecture, achieving accuracy comparable to top-down methods while maintaining high speed. We propose a dynamic coordi-nate classifier and a tailored loss function for heatmap learning, specifically designed to address the incompati-bilities between coordinate classification and dense pre-diction models. RTMO outperforms state-of-the-art one-stage pose estimators, achieving 1.1% higher AP on COCO while operating about 9 times faster with the same back-bone. Our largest model, RTMO-1, attains 74.8% AP on COCO va12017 and 141 FPS on a single V100 GPU, demonstrating its efficiency and accuracy. The code and models are available at https://github.com/open-mmlab/mmpose/tree/main/projects/rtmo.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Referring Human Pose and Mask Estimation In the WildBo Miao, Mingtao Feng, Zijie Wu, Mohammed Bennamoun et al.NeurIPS 2024 · 12 citations
- EDM: Efficient Deep Feature MatchingXi Li, Tong Rao, Cihui PanICCV 2025 · 6 citations
- FIP: Endowing Robust Motion Capture on Daily Garment by Fusing Flex and Inertial SensorsRuonan Zheng, Jiawei Fang, Yuan Yao, Xiaoxia Gao et al.CHI 2025 · 5 citations
- MultiCam: On-the-fly Multi-Camera Pose Estimation Using Spatiotemporal Overlaps of Known ObjectsShiyu Li, Hannah Schieber, Kristoffer Waldow, Benjamin Busam et al.IEEE VR 2026 · 2 citations
- End-to-End Multi-Person Pose Estimation with Pose-Aware Video TransformerYonghui Yu, Jiahang Cai, Xun Wang, Wenwu YangAAAI 2026 · 2 citations
Builds on19
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi et al.ICCV 2019 · 3,348 citations
- DETRs Beat YOLOs on Real-time Object DetectionYian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei et al.CVPR 2024 · 3,046 citations
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 1,105 citations
Related papers
- InsPose: Instance-Aware Networks for Single-Stage Multi-Person Pose EstimationDahu Shi, Xing Wei, Xiaodong Yu, Wenming Tan et al.ACM MM 2021 · 40 citations
- FCPose: Fully Convolutional Multi-Person Pose Estimation With Dynamic Instance-Aware ConvolutionsWeian Mao, Zhi Tian, Xinlong Wang, Chunhua ShenCVPR 2021
- Single-Network Whole-Body Pose EstimationGines Hidalgo Martinez, Yaadhav Raaj, Haroon Idrees, Donglai Xiang et al.ICCV 2019 · 115 citations
- SIMPLE: SIngle-network with Mimicking and Point Learning for Bottom-up Human Pose EstimationJiabin Zhang, Zheng Zhu, Jiwen Lu, Junjie Huang et al.AAAI 2021 · 14 citations
- Learning Local-Global Contextual Adaptation for Multi-Person Pose EstimationNan Xue, Tianfu Wu, Gui-Song Xia, Liangpei ZhangCVPR 2022 · 42 citations
