PoseIRM: Enhance 3D Human Pose Estimation on Unseen Camera Settings via Invariant Risk Minimization
Yanlu Cai, Weizhong Zhang, Yuan Wu, Cheng Jin
摘要
Camera-parameter-free multi-view pose estimation is an emerging technique for 3D human pose estimation (HPE). They can infer the camera settings implicitly or explicitly to mitigate the depth uncertainty impact, showcasing significant potential in real applications. However, due to the limited camera setting diversity in the available datasets, the inferred camera parameters are always simply hardcoded into the model during training and not adaptable to the input in inference, making the learned models cannot generalize well under unseen camera settings. A natural solution is to artificially synthesize some samples, i.e., 2D-3D pose pairs, under massive new camera settings. Unfortunately, to prevent over-fitting the existing camera setting, the number of synthesized samples for each new camera setting should be comparable with that for the existing one, which multiplies the scale of training and even makes it computationally prohibitive. In this paper, we propose a novel HPE approach under the invariant risk minimization (IRM) paradigm. Precisely, we first synthesize 2D poses from myriad camera settings. We then train our model under the IRM paradigm, which targets at learning a common optimal model across all camera settings and thus enforces the model to automatically learn the camera parameters based on the input data. This allows the model to accurately infer 3D poses on unseen data by training on only a handful of samples from each synthesized setting and thus avoid the unbearable training cost increment. Another appealing feature of our method is that benefited from the capability of IRM in identifying the invariant features, its performance on the seen camera settings is enhanced as well. Comprehensive experiments verify the superiority of our approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- PrimeComposer: Faster Progressively Combined Diffusion for Image Composition with Attention SteeringYibin Wang, Weizhong Zhang, Jianwei Zheng, Cheng JinACM MM 2024 · 被引用 9 次
- Unified 2D-3D Discrete Priors for Noise-Robust and Calibration-Free Multiview 3D Human Pose EstimationGeng Chen, Pengfei Ren, Xufeng Jian, Haifeng Sun 等NeurIPS 2025 · 被引用 1 次
它引用的顶会 Paper13
- 3D Human Pose Estimation with Spatial and Temporal TransformersCe Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang 等ICCV 2021 · 被引用 648 次
- Learnable Triangulation of Human PoseKarim Iskakov, Egor Burkov, Victor S. Lempitsky, Yury MalkovICCV 2019 · 被引用 419 次
- MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in VideoJinlu Zhang, Zhigang Tu, Jianyu Yang, Yujin Chen 等CVPR 2022 · 被引用 356 次
- Cross View Fusion for 3D Human Pose EstimationHaibo Qiu, Chunyu Wang, Jingdong Wang, Naiyan Wang 等ICCV 2019 · 被引用 242 次
- Invariant RationalizationShiyu Chang, Yang Zhang, Mo Yu, Tommi S. JaakkolaICML 2020 · 被引用 232 次
相关 Paper
- Weakly-Supervised 3D Human Pose Learning via Multi-View Images in the WildUmar Iqbal, Pavlo Molchanov, Jan KautzCVPR 2020
- FusionFormer: A Concise Unified Feature Fusion Transformer for 3D Pose EstimationYanlu Cai, Weizhong Zhang, Yuan Wu, Cheng JinAAAI 2024 · 被引用 23 次
- AdaptPose: Cross-Dataset Adaptation for 3D Human Pose Estimation by Learnable Motion GenerationMohsen Gholami, Bastian Wandt, Helge Rhodin, Rabab Ward 等CVPR 2022 · 被引用 31 次
- CanonPose: Self-Supervised Monocular 3D Human Pose Estimation in the WildBastian Wandt, Marco Rudolph, Petrissa Zell, Helge Rhodin 等CVPR 2021
- Bayesian Invariant Risk MinimizationYong Lin, Hanze Dong, Hao Wang, Tong ZhangCVPR 2022 · 被引用 48 次
