DECA: Deep viewpoint-Equivariant human pose estimation using Capsule Autoencoders
Nicola Garau, Niccolò Bisagno, Piotr Bródka, Nicola Conci
摘要
Human Pose Estimation (HPE) aims at retrieving the 3D position of human joints from images or videos. We show that current 3D HPE methods suffer a lack of viewpoint equivariance, namely they tend to fail or perform poorly when dealing with viewpoints unseen at training time. Deep learning methods often rely on either scale-invariant, translation-invariant, or rotation-invariant operations, such as max-pooling. However, the adoption of such procedures does not necessarily improve viewpoint generalization, rather leading to more data-dependent methods. To tackle this issue, we propose a novel capsule autoencoder network with fast Variational Bayes capsule routing, named DECA. By modeling each joint as a capsule entity, combined with the routing algorithm, our approach can preserve the joints’ hierarchical and geometrical structure in the feature space, independently from the viewpoint. By achieving viewpoint equivariance, we drastically reduce the network data dependency at training time, resulting in an improved ability to generalize for unseen viewpoints. In the experimental validation, we outperform other methods on depth images from both seen and unseen viewpoints, both top-view, and front-view. In the RGB domain, the same network gives state-of-the-art results on the challenging viewpoint transfer task, also establishing a new framework for top-view HPE. The code can be found at https://github.com/mmlab-cv/DECA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- GazeOnce: Real-Time Multi-Person Gaze EstimationMingfang Zhang, Yunfei Liu, Feng LuCVPR 2022 · 被引用 30 次
- Lightweight Super-Resolution Head for Human Pose EstimationHaonan Wang, Jie Liu, Jie Tang, Gangshan WuACM MM 2023 · 被引用 23 次
它引用的顶会 Paper4
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 被引用 1,139 次
- A2J: Anchor-to-Joint Regression Network for 3D Articulated Pose Estimation From a Single Depth ImageFu Xiong, Boshen Zhang, Yang Xiao, Zhiguo Cao 等ICCV 2019 · 被引用 178 次
- Capsule Routing via Variational BayesFabio De Sousa Ribeiro, Georgios Leontidis, Stefanos D. KolliasAAAI 2020 · 被引用 93 次
- VIBE: Video Inference for Human Body Pose and Shape EstimationMuhammed Kocabas, Nikos Athanasiou, Michael J. BlackCVPR 2020
相关 Paper
- Equicaps: Predictor-Free Pose-Aware Pre-Trained Capsule NetworksAthinoulla Konstantinou, Georgios Leontidis, Mamatha Thota, Aiden DurrantICCV 2025
- Building Deep Equivariant Capsule NetworksSai Raam Venkataraman, S. Balasubramanian, R. Raghunatha SarmaICLR 2020 · 被引用 28 次
- Unsupervised Motion Representation Learning with Capsule AutoencodersZiwei Xu, Xudong Shen, Yongkang Wong, Mohan S. KankanhalliNeurIPS 2021 · 被引用 33 次
- Monocular 3D Human Pose Estimation by Generation and Ordinal RankingSaurabh Sharma, Pavan Teja Varigonda, Prashast Bindal, Abhishek Sharma 等ICCV 2019 · 被引用 177 次
- GenCape: Structure-Inductive Generative Modeling for Category-Agnostic Pose EstimationJiyong Rao, Yu Wang, Shengjie ZhaoICLR 2026
