CapeNext: Rethinking and Refining Dynamic Support Information for Category-Agnostic Pose Estimation
Yu Zhu, Dan Zeng, Shuiwang Li, Qijun Zhao, Qiaomu Shen, Bo Tang
Abstract
Recent research in Category-Agnostic Pose Estimation (CAPE) has adopted fixed textual keypoint description as semantic prior for two-stage pose matching frameworks. While this paradigm enhances robustness and flexibility by disentangling the dependency of support images, our critical analysis reveals two inherent limitations of static joint embedding: (1) polysemy-induced cross-category ambiguity during the matching process(e.g., the concept "leg" exhibiting divergent visual manifestations across humans and furniture), and (2) insufficient discriminability for fine-grained intra-category variations (e.g., posture and fur discrepancies between a sleeping white cat and a standing black cat). To overcome these challenges, we propose a new framework that innovatively integrates hierarchical cross-modal interaction with dual-stream feature refinement, enhancing the joint embedding with both class-level and instance-specific cues from textual description and specific images. Experiments on the MP-100 dataset demonstrate that, regardless of the network backbone, CapeNext consistently outperforms state-of-the-art CAPE methods by a large margin.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- TokenPose: Learning Keypoint Tokens for Human Pose EstimationYanjie Li, Shoukui Zhang, Zhicheng Wang, Sen Yang et al.ICCV 2021 · 363 citations
- A Simple Framework for Open-Vocabulary Segmentation and DetectionHao Zhang, Feng Li, Xueyan Zou, Shilong Liu et al.ICCV 2023 · 241 citations
Related papers
- Meta-Point Learning and Refining for Category-Agnostic Pose EstimationJunjie Chen, Jiebin Yan, Yuming Fang, Li NiuCVPR 2024
- CapeX: Category-Agnostic Pose Estimation from Textual Point ExplanationMatan Rusanovsky, Or Hirschorn, Shai AvidanICLR 2025
- CapeLLM: Support-Free Category-Agnostic Pose Estimation with Multimodal Large Language ModelsJunho Kim, Hyungjin Chung, Byung-Hoon KimICCV 2025 · 1 citation
- GenCape: Structure-Inductive Generative Modeling for Category-Agnostic Pose EstimationJiyong Rao, Yu Wang, Shengjie ZhaoICLR 2026
- Dynamic Support Information Mining for Category-Agnostic Pose EstimationPengfei Ren, Yuanyuan Gao, Haifeng Sun, Qi Qi et al.CVPR 2024 · 3 citations
