Matching Is Not Enough: A Two-Stage Framework for Category-Agnostic Pose Estimation
Min Shi, Zihao Huang, Xianzheng Ma, Xiaowei Hu, Zhiguo Cao
Abstract
Category-agnostic pose estimation (CAPE) aims to predict keypoints for arbitrary categories given support images with keypoint annotations. Existing approaches match the keypoints across the image for localization. However, such a one-stage matching paradigm shows inferior accuracy: the prediction heavily relies on the matching results, which can be noisy due to the open set nature in CAPE. For example, two mirror-symmetric keypoints (e.g., left and right eyes) in the query image can both trigger high similarity on certain support keypoints (eyes), which leads to duplicated or opposite predictions. To calibrate the inaccurate matching results, we introduce a two-stage framework, where matched keypoints from the first stage are viewed as similarity-aware position proposals. Then, the model learns to fetch relevant features to correct the initial proposals in the second stage. We instantiate the framework with a transformer model tailored for CAPE. The transformer encoder incorporates specific designs to improve the representation and similarity modeling in the first matching stage. In the second stage, similarity-aware proposals are packed as queries in the decoder for refinement via cross-attention. Our method surpasses the previous best approach by large margins on CAPE benchmark MP-100 on both accuracy and efficiency. Code available at github.com/flyinglynx/CapeFormer
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0bf92dec-5455-4dfc-83c7-0dab5488207aCited by top-tier papers15
- Detect Any Keypoints: An Efficient Light-Weight Few-Shot Keypoint DetectorChangsheng Lu, Piotr KoniuszAAAI 2024 · 12 citations
- KptLLM: Unveiling the Power of Large Language Model for Keypoint ComprehensionJie Yang, Wang Zeng, Sheng Jin, Lumin Xu et al.NeurIPS 2024 · 9 citations
- MoCapAnything: Unified 3D Motion Capture for Arbitrary Skeletons from Monocular VideosKehong Gong, Zhengyu Wen, Xiaoyu He, Mingxi Xu et al.CVPR 2026 · 8 citations
- Weak-shot Keypoint Estimation via Keyness and Correspondence TransferJunjie Chen, Zeyu Luo, Zezheng Liu, Wenhui Jiang et al.NeurIPS 2025 · 5 citations
- Dynamic Support Information Mining for Category-Agnostic Pose EstimationPengfei Ren, Yuanyuan Gao, Haifeng Sun, Qi Qi et al.CVPR 2024 · 3 citations
Builds on12
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng et al.ICCV 2021 · 974 citations
- Human Pose Regression with Residual Log-likelihood EstimationJiefeng Li, Siyuan Bian, Ailing Zeng, Can Wang et al.ICCV 2021 · 286 citations
- Single-Stage Multi-Person Pose MachinesXuecheng Nie, Jiashi Feng, Jianfeng Zhang, Shuicheng YanICCV 2019 · 246 citations
- Cross-Domain Adaptation for Animal Pose EstimationJinkun Cao, Hongyang Tang, Haoshu Fang, Xiaoyong Shen et al.ICCV 2019 · 209 citations
Related papers
- Meta-Point Learning and Refining for Category-Agnostic Pose EstimationJunjie Chen, Jiebin Yan, Yuming Fang, Li NiuCVPR 2024
- GenCape: Structure-Inductive Generative Modeling for Category-Agnostic Pose EstimationJiyong Rao, Yu Wang, Shengjie ZhaoICLR 2026
- EdgeCape: Edge Weight Prediction For Category-Agnostic Pose EstimationOr Hirschorn, Shai AvidanICLR 2026 · 1 citation
- ESCAPE: Encoding Super-keypoints for Category-Agnostic Pose EstimationKhoi Duc Nguyen, Chen Li, Gim Hee LeeCVPR 2024 · 1 citation
- CapeX: Category-Agnostic Pose Estimation from Textual Point ExplanationMatan Rusanovsky, Or Hirschorn, Shai AvidanICLR 2025
