POTTER: Pooling Attention Transformer for Efficient Human Mesh Recovery
Ce Zheng, Xianpeng Liu, Guo-Jun Qi, Chen Chen
摘要
Transformer architectures have achieved SOTA performance on the human mesh recovery (HMR) from monocular images. However, the performance gain has come at the cost of substantial memory and computational overhead. A lightweight and efficient model to reconstruct accurate human mesh is needed for real-world applications. In this paper, we propose a pure transformer architecture named POoling aTtention TransformER (POTTER) for the HMR task from single images. Observing that the conventional attention module is memory and computationally expensive, we propose an efficient pooling attention module, which significantly reduces the memory and computational cost without sacrificing performance. Furthermore, we design a new transformer architecture by integrating a High-Resolution (HR) stream for the HMR task. The high-resolution local and global features from the HR stream can be utilized for recovering more accurate human mesh. Our POTTER outperforms the SOTA method METRO by only requiring 7% of total parameters and 14% of the Multiply-Accumulate Operations on the Human3.6M (PA-MPJPE metric) and 3DPW (all three metrics) datasets. The project webpage is https://zczcwh.github.io/potter_page/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- A Single 2D Pose with Context is Worth Hundreds for 3D Human Pose EstimationQitao Zhao, Ce Zheng, Mengyuan Liu, Chen ChenNeurIPS 2023 · 被引用 40 次
- A Dual-Augmentor Framework for Domain Generalization in 3D Human Pose EstimationQucheng Peng, Ce Zheng, Chen ChenCVPR 2024 · 被引用 38 次
- PostureHMR: Posture Transformation for 3D Human Mesh RecoveryYu-Pei Song, Xiao Wu, Zhaoquan Yuanl, Jian-Jun Qiao 等CVPR 2024 · 被引用 11 次
- ScoreHypo: Probabilistic Human Mesh Estimation with Hypothesis ScoringYuan Xu, Xiaoxuan Ma, Jiajun Su, Wentao Zhu 等CVPR 2024 · 被引用 6 次
- Toward Approaches to Scalability in 3D Human Pose EstimationJun-Hui Kim, Seong-Whan LeeNeurIPS 2024 · 被引用 5 次
它引用的顶会 Paper27
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
相关 Paper
- Deformable Mesh Transformer for 3D Human Mesh RecoveryYusuke YoshiyasuCVPR 2023
- End-to-End Human Pose and Mesh Reconstruction with TransformersKevin Lin, Lijuan Wang, Zicheng LiuCVPR 2021
- FeatER: An Efficient Network for Human Reconstruction via Feature Map-Based TransformERCe Zheng, Matías Mendieta, Taojiannan Yang, Guo-Jun Qi 等CVPR 2023
- TORE: Token Reduction for Efficient Human Mesh Recovery with TransformerZhiyang Dou, Qingxuan Wu, Cheng Lin, Zeyu Cao 等ICCV 2023 · 被引用 56 次
- Mesh GraphormerKevin Lin, Lijuan Wang, Zicheng LiuICCV 2021 · 被引用 399 次
