CIGPose: Causal Intervention Graph Neural Network for Whole-Body Pose Estimation
Bohao Li, Zhicheng Cao, Huixian Li, Yangming Guo
摘要
State-of-the-art whole-body pose estimators often lack robustness, producing anatomically implausible predictions in challenging scenes. We posit this failure stems from spurious correlations learned from visual context, a problem we formalize using a Structural Causal Model (SCM). The SCM identifies visual context as a confounder that creates a non-causal backdoor path, corrupting the model's reasoning. We introduce the Causal Intervention Graph Pose (CIGPose) framework to address this by approximating the true causal effect between visual evidence and pose. The core of CIGPose is a novel Causal Intervention Module: it first identifies confounded keypoint representations via predictive uncertainty and then replaces them with learned, context-invariant canonical embeddings. These deconfounded embeddings are processed by a hierarchical graph neural network that reasons over the human skeleton at both local and global semantic levels to enforce anatomical plausibility. Extensive experiments show CIGPose achieves a new state-of-the-art on COCO-WholeBody. Notably, our CIGPose-x model achieves 67.0% AP, surpassing prior methods that rely on extra training data. With the additional UBody dataset, CIGPose-x is further boosted to 67.5% AP, demonstrating superior robustness and data efficiency. The codes and models are publicly available at https://github.com/53mins/CIGPose.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper30
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 被引用 1,105 次
- Causal Intervention for Weakly-Supervised Semantic SegmentationDong Zhang, Hanwang Zhang, Jinhui Tang, Xian-Sheng Hua 等NeurIPS 2020 · 被引用 563 次
- Learning Graph Convolutional Network for Skeleton-Based Human Action Recognition by Neural SearchingWei Peng, Xiaopeng Hong, Haoyu Chen, Guoying ZhaoAAAI 2020 · 被引用 362 次
- TransPose: Keypoint Localization via TransformerSen Yang, Zhibin Quan, Mu Nie, Wankou YangICCV 2021 · 被引用 360 次
相关 Paper
- Context Modeling in 3D Human Pose Estimation: A Unified PerspectiveXiaoxuan Ma, Jiajun Su, Chunyu Wang, Hai Ci 等CVPR 2021
- CleanPose: Category-Level Object Pose Estimation via Causal Learning and Knowledge DistillationXiao Lin, Yun Peng, Liuyi Wang, Xianyou Zhong 等ICCV 2025 · 被引用 3 次
- Contextual Debiasing for Visual Recognition with Causal MechanismsRuyang Liu, Hao Liu, Ge Li, Haodi Hou 等CVPR 2022 · 被引用 42 次
- Adaptive Hypergraph Neural Network for Multi-Person Pose EstimationXixia Xu, Qi Zou, Xue LinAAAI 2022 · 被引用 14 次
- Causal-Inspired Multitask Learning for Video-Based Human Pose EstimationHaipeng Chen, Sifan Wu, Zhigang Wang, Yifang Yin 等AAAI 2025 · 被引用 7 次
