CapeX: Category-Agnostic Pose Estimation from Textual Point Explanation
Matan Rusanovsky, Or Hirschorn, Shai Avidan
Abstract
Conventional 2D pose estimation models are constrained by their design to specific object categories. This limits their applicability to predefined objects. To overcome these limitations, category-agnostic pose estimation (CAPE) emerged as a solution. CAPE aims to facilitate keypoint localization for diverse object categories using a unified model, which can generalize from minimal annotated support images. Recent CAPE works have produced object poses based on arbitrary keypoint definitions annotated on a user-provided support image. Our work departs from conventional CAPE methods, which require a support image, by adopting a text-based approach instead of the support image. Specifically, we use a pose-graph, where nodes represent keypoints that are described with text. This representation takes advantage of the abstraction of text descriptions and the structure imposed by the graph. Our approach effectively breaks symmetry, preserves structure, and improves occlusion handling. We validate our novel approach using the MP-100 benchmark, a comprehensive dataset covering over 100 categories and 18,000 images. MP-100 is structured so that the evaluation categories are unseen during training, making it especially suited for CAPE. Under a 1-shot setting, our solution achieves a notable performance boost of 1.26%, establishing a new state-of-the-art for CAPE. Additionally, we enhance the dataset by providing text description annotations for both training and testing. We also include alternative text annotations specifically for testing the model's ability to generalize across different textual descriptions, further increasing its value for future research. Our code and dataset are publicly available at https://github.com/matanr/capex .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e9292ef5-d2d8-435d-a1ea-f6fef2dfeaa1Cited by top-tier papers8
- MoCapAnything: Unified 3D Motion Capture for Arbitrary Skeletons from Monocular VideosKehong Gong, Zhengyu Wen, Xiaoyu He, Mingxi Xu et al.CVPR 2026 · 8 citations
- Weak-shot Keypoint Estimation via Keyness and Correspondence TransferJunjie Chen, Zeyu Luo, Zezheng Liu, Wenhui Jiang et al.NeurIPS 2025 · 5 citations
- CapeLLM: Support-Free Category-Agnostic Pose Estimation with Multimodal Large Language ModelsJunho Kim, Hyungjin Chung, Byung-Hoon KimICCV 2025 · 1 citation
- EdgeCape: Edge Weight Prediction For Category-Agnostic Pose EstimationOr Hirschorn, Shai AvidanICLR 2026 · 1 citation
- Recurrent Feature Mining and Keypoint Mixup Padding for Category-Agnostic Pose EstimationJunjie Chen, Weilong Chen, Yifan Zuo, Yuming FangCVPR 2025
Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
- TransPose: Keypoint Localization via TransformerSen Yang, Zhibin Quan, Mu Nie, Wankou YangICCV 2021 · 360 citations
- Represent, Compare, and Learn: A Similarity-Aware Framework for Class-Agnostic CountingMin Shi, Hao Lu, Chen Feng, Chengxin Liu et al.CVPR 2022 · 99 citations
- Dynamic Support Information Mining for Category-Agnostic Pose EstimationPengfei Ren, Yuanyuan Gao, Haifeng Sun, Qi Qi et al.CVPR 2024 · 3 citations
Related papers
- GenCape: Structure-Inductive Generative Modeling for Category-Agnostic Pose EstimationJiyong Rao, Yu Wang, Shengjie ZhaoICLR 2026
- ESCAPE: Encoding Super-keypoints for Category-Agnostic Pose EstimationKhoi Duc Nguyen, Chen Li, Gim Hee LeeCVPR 2024 · 1 citation
- Meta-Point Learning and Refining for Category-Agnostic Pose EstimationJunjie Chen, Jiebin Yan, Yuming Fang, Li NiuCVPR 2024
- CapeNext: Rethinking and Refining Dynamic Support Information for Category-Agnostic Pose EstimationYu Zhu, Dan Zeng, Shuiwang Li, Qijun Zhao et al.AAAI 2026
- Matching Is Not Enough: A Two-Stage Framework for Category-Agnostic Pose EstimationMin Shi, Zihao Huang, Xianzheng Ma, Xiaowei Hu et al.CVPR 2023
