PCLs: Geometry-Aware Neural Reconstruction of 3D Pose With Perspective Crop Layers
Frank Yu, Mathieu Salzmann, Pascal Fua, Helge Rhodin
摘要
Local processing is an essential feature of CNNs and other neural network architectures-it is one of the reasons why they work so well on images where relevant information is, to a large extent, local. However, perspective effects stemming from the projection in a conventional camera vary for different global positions in the image. We introduce Perspective Crop Layers (PCLs)-a form of perspective crop of the region of interest based on the camera geometry-and show that accounting for the perspective consistently improves the accuracy of state-of-theart 3D pose reconstruction methods. PCLs are modular neural network layers, which, when inserted into existing CNN and MLP architectures, deterministically remove the location-dependent perspective effects while leaving end-to-end training and the number of parameters of the underlying neural network unchanged. We demonstrate that PCL leads to improved 3D human pose reconstruction accuracy for CNN architectures that use cropping operations, such as spatial transformer networks (STN), and, somewhat surprisingly, MLPs used for 2D-to-3D keypoint lifting. Our conclusion is that it is important to utilize camera calibration information when available, for classical and deep-learning-based computer vision alike. PCL offers an easy way to improve the accuracy of existing 3D reconstruction networks by making them geometryaware. Our code is publicly available at github.com/yu- frank/PerspectiveCropLayers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- SPEC: Seeing People in the Wild with an Estimated CameraMuhammed Kocabas, Chun-Hao P. Huang, Joachim Tesch, Lea Müller 等ICCV 2021 · 被引用 181 次
- AdaptPose: Cross-Dataset Adaptation for 3D Human Pose Estimation by Learnable Motion GenerationMohsen Gholami, Bastian Wandt, Helge Rhodin, Rabab Ward 等CVPR 2022 · 被引用 31 次
- HDG-ODE: A Hierarchical Continuous-Time Model for Human Pose ForecastingYucheng Xing, Xin WangICCV 2023 · 被引用 5 次
- HORT: Monocular Hand-held Objects Reconstruction with TransformersZerui Chen, Rolandos Alexandros Potamias, Shizhe Chen, Cordelia SchmidICCV 2025 · 被引用 4 次
- A Light Touch Approach to Teaching Transformers Multi-view GeometryYash Bhalgat, João F. Henriques, Andrew ZissermanCVPR 2023
它引用的顶会 Paper5
- XNect: real-time multi-person 3D motion capture with a single RGB cameraDushyant Mehta, Oleksandr Sotnychenko, Franziska Mueller, Weipeng Xu 等SIGGRAPH 2020 · 被引用 267 次
- HEMlets Pose: Learning Part-Centric Heatmap Triplets for Accurate 3D Human Pose EstimationKun Zhou, Xiaoguang Han, Nianjuan Jiang, Kui Jia 等ICCV 2019 · 被引用 129 次
- Learning Perspective Undistortion of PortraitsYajie Zhao, Zeng Huang, Tianye Li, Weikai Chen 等ICCV 2019 · 被引用 29 次
- Cascaded Deep Monocular 3D Human Pose Estimation With Evolutionary Training DataShichao Li, Lei Ke, Kevin Pratama, Yu-Wing Tai 等CVPR 2020
- Deep Kinematics Analysis for Monocular 3D Human Pose EstimationJingwei Xu, Zhenbo Yu, Bingbing Ni, Jiancheng Yang 等CVPR 2020
相关 Paper
- 3D-LFM: Lifting Foundation ModelMosam Dabhi, László A. Jeni, Simon LuceyCVPR 2024
- Learning Viewpoint-Agnostic Visual Representations by Recovering Tokens in 3D SpaceJinghuan Shang, Srijan Das, Michael S. RyooNeurIPS 2022 · 被引用 18 次
- Lightweight Multi-View 3D Pose Estimation Through Camera-Disentangled RepresentationEdoardo Remelli, Shangchen Han, Sina Honari, Pascal Fua 等CVPR 2020
- Why Having 10, 000 Parameters in Your Camera Model Is Better Than TwelveThomas Schöps, Viktor Larsson, Marc Pollefeys, Torsten SattlerCVPR 2020
- LitePT: Lighter Yet Stronger Point TransformerYuanwen Yue, Damien Robert, Jianyuan Wang, Sunghwan Hong 等CVPR 2026 · 被引用 25 次
