Group Pose: A Simple Baseline for End-to-End Multi-person Pose Estimation
Huan Liu, Qiang Chen, Zichang Tan, Jiang-Jiang Liu, Jian Wang, Xiangbo Su, Xiaolong Li, Kun Yao, Junyu Han, Errui Ding, Yao Zhao, Jingdong Wang
摘要
In this paper, we study the problem of end-to-end multi-person pose estimation. State-of-the-art solutions adopt the DETR-like framework, and mainly develop the complex decoder, e.g., regarding pose estimation as keypoint box detection and combining with human detection in ED-Pose [38], hierarchically predicting with pose decoder and joint (keypoint) decoder in PETR [27].We present a simple yet effective transformer approach, named Group Pose. We simply regard K-keypoint pose estimation as predicting a set of N × K keypoint positions, each from a keypoint query, as well as representing each pose with an instance query for scoring N pose predictions.Motivated by the intuition that the interaction, among across-instance queries of different types, is not directly helpful, we make a simple modification to decoder self-attention. We replace single self-attention over all the N × (K + 1) queries with two subsequent group self-attentions: (i) N within-instance self-attention, with each over K keypoint queries and one instance query, and (ii) (K +1) same-type across-instance self-attention, each over N queries of the same type. The resulting decoder removes the interaction among across-instance type-different queries, easing the optimization and thus improving the performance. Experimental results on MS COCO and Crowd-Pose show that our approach without human box supervision is superior to previous methods with complex decoders, and even is slightly better than ED-Pose that uses human box supervision. Paddle 1 and PyTorch 2 codes are available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language TasksJiannan Wu, Muyan Zhong, Sen Xing, Zeqiang Lai 等NeurIPS 2024 · 被引用 179 次
- DV-3DLane: End-to-end Multi-modal 3D Lane Detection with Dual-view RepresentationYueru Luo, Shuguang Cui, Zhen LiICLR 2024 · 被引用 15 次
- Referring Human Pose and Mask Estimation In the WildBo Miao, Mingtao Feng, Zijie Wu, Mohammed Bennamoun 等NeurIPS 2024 · 被引用 12 次
- DiffusionRegPose: Enhancing Multi-Person Pose Estimation Using a Diffusion-Based End-to-End Regression ApproachDayi Tan, Hansheng Chen, Wei Tian, Lu XiongCVPR 2024 · 被引用 6 次
- Weak-shot Keypoint Estimation via Keyness and Correspondence TransferJunjie Chen, Zeyu Luo, Zezheng Liu, Wenhui Jiang 等NeurIPS 2025 · 被引用 5 次
它引用的顶会 Paper17
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang 等ICLR 2022 · 被引用 1,218 次
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng 等ICCV 2021 · 被引用 974 次
- DN-DETR: Accelerate DETR Training by Introducing Query DeNoisingFeng Li, Hao Zhang, Shilong Liu, Jian Guo 等CVPR 2022 · 被引用 879 次
相关 Paper
- End-to-End Multi-Person Pose Estimation with TransformersDahu Shi, Xing Wei, Liangqi Li, Ye Ren 等CVPR 2022 · 被引用 147 次
- The Center of Attention: Center-Keypoint Grouping via Attention for Multi-Person Pose EstimationGuillem Brasó, Nikita Kister, Laura Leal-TaixéICCV 2021 · 被引用 50 次
- Bottom-Up Human Pose Estimation via Disentangled Keypoint RegressionZigang Geng, Ke Sun, Bin Xiao, Zhaoxiang Zhang 等CVPR 2021
- PSVT: End-to-End Multi-Person 3D Pose and Shape Estimation with Progressive Video TransformersZhongwei Qiu, Qiansheng Yang, Jian Wang, Haocheng Feng 等CVPR 2023
- End-to-End Multi-Person Pose Estimation with Pose-Aware Video TransformerYonghui Yu, Jiahang Cai, Xun Wang, Wenwu YangAAAI 2026 · 被引用 2 次
