Rethinking pose estimation in crowds: overcoming the detection information bottleneck and ambiguity
Mu Zhou, Lucas Stoffl, Mackenzie Weygandt Mathis, Alexander Mathis
Abstract
Frequent interactions between individuals are a fundamental challenge for pose estimation algorithms. Current pipelines either use an object detector together with a pose estimator (top-down approach), or localize all body parts first and then link them to predict the pose of individuals (bottom-up). Yet, when individuals closely interact, top-down methods are ill-defined due to overlapping individuals, and bottom-up methods often falsely infer connections to distant bodyparts. Thus, we propose a novel pipeline called bottom-up conditioned top-down pose estimation (BUCTD) that combines the strengths of bottomup and top-down methods. Specifically, we propose to use a bottom-up model as the detector, which in addition to an estimated bounding box provides a pose proposal that is fed as condition to an attention-based top-down model. We demonstrate the performance and efficiency of our approach on animal and human pose estimation benchmarks. On CrowdPose and OCHuman, we outperform previous state-of-the-art models by a significant margin. We achieve 78.5 AP on CrowdPose and 48.5 AP on OCHuman, an improvement of 8.6% and 7.8% over the prior art, respectively. Furthermore, we show that our method strongly improves the performance on multi-animal benchmarks involving fish and monkeys. The code is available at https://github.com/amathislab/BUCTD
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5c62c1f0-3410-4d7d-8b6e-edfdd00463adCited by top-tier papers4
- Seeing Beyond the Crop: Using Language Priors for Out-of-Bounding Box Keypoint PredictionBavesh Balaji, Jerrin Bright, Yuhao Chen, Sirisha Rambhatla et al.NeurIPS 2024 · 4 citations
- Detection, Pose Estimation and Segmentation for Multiple Bodies: Closing the Virtuous CircleMiroslav Purkrábek, Jiri MatasICCV 2025
- Multi-Agent Long-Term 3D Human Pose Forecasting via Interaction-Aware Trajectory ConditioningJaewoo Jeong, Daehee Park, Kuk-Jin YoonCVPR 2024
- Adversarially Robust Out-of-Distribution Detection Using Lyapunov-Stabilized EmbeddingsHossein Mirzaei, Mackenzie W. MathisICLR 2025
Builds on13
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 1,105 citations
- Symbolic Discovery of Optimization AlgorithmsXiangning Chen, Chen Liang, Da Huang, Esteban Real et al.NeurIPS 2023 · 734 citations
- Exploiting Spatial-Temporal Relationships for 3D Pose Estimation via Graph Convolutional NetworksYujun Cai, Liuhao Ge, Jun Liu, Jianfei Cai et al.ICCV 2019 · 504 citations
- TokenPose: Learning Keypoint Tokens for Human Pose EstimationYanjie Li, Shoukui Zhang, Zhicheng Wang, Sen Yang et al.ICCV 2021 · 363 citations
Related papers
- Occluded Human Mesh RecoveryRawal Khirodkar, Shashank Tripathi, Kris KitaniCVPR 2022 · 74 citations
- Multi-Instance Pose Networks: Rethinking Top-Down Pose EstimationRawal Khirodkar, Visesh Chari, Amit Agrawal, Ambrish TyagiICCV 2021 · 80 citations
- A Characteristic Function-Based Method for Bottom-Up Human Pose EstimationHaoxuan Qu, Yujun Cai, Lin Geng Foo, Ajay Kumar et al.CVPR 2023
- Combining Detection and Tracking for Human Pose Estimation in VideosManchen Wang, Joseph Tighe, Davide ModoloCVPR 2020
- SIMPLE: SIngle-network with Mimicking and Point Learning for Bottom-up Human Pose EstimationJiabin Zhang, Zheng Zhu, Jiwen Lu, Junjie Huang et al.AAAI 2021 · 14 citations
