HOT-Net: Non-Autoregressive Transformer for 3D Hand-Object Pose Estimation
Lin Huang, Jianchao Tan, Jingjing Meng, Ji Liu, Junsong Yuan
摘要
As we use our hands frequently in daily activities, the analysis of hand-object interactions plays a critical role to many multimedia understanding and interaction applications. Different from conventional 3D hand-only and object-only pose estimation, estimating 3D hand-object pose is more challenging due to the mutual occlusions between hand and object, as well as the physical constraints between them. To overcome these issues, we propose to fully utilize the structural correlations among hand joints and object corners in order to obtain more reliable poses. Our work is inspired by structured output learning models in sequence transduction field like Transformer encoder-decoder framework. Besides modeling inherent dependencies from extracted 2D hand-object pose, our proposed Hand-Object Transformer Network (HOT-Net) also captures the structural correlations among 3D hand joints and object corners. Similar to Transformer's autoregressive decoder, by considering structured output patterns, this helps better constrain the output space and leads to more robust pose estimation. However, different from Transformer's sequential modeling mechanism, HOT-Net adopts a novel non-autoregressive decoding strategy for 3D hand-object pose estimation. Specifically, our model removes the Transformer's dependence on previously generated results and explicitly feeds a reference 3D hand-object pose into the decoding process to provide equivalent target pose patterns for parallely localizing each 3D keypoint. To further improve physical validity of estimated hand pose, besides anatomical constraints, we propose a cooperative pose constraint, aiming to enable the hand pose to cooperate with hand shape, to generate hand mesh. We demonstrate real-time speed and state-of-the-art performance on benchmark hand-object datasets for both 3D hand and object poses.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Not All Tokens Are Equal: Human-centric Visual Analysis via Token Clustering TransformerWang Zeng, Sheng Jin, Wentao Liu, Chen Qian 等CVPR 2022 · 被引用 132 次
- DeepMultiCap: Performance Capture of Multiple Characters Using Sparse Multiview CamerasYang Zheng, Ruizhi Shao, Yuxiang Zhang, Tao Yu 等ICCV 2021 · 被引用 112 次
- E2VPT: An Effective and Efficient Approach for Visual Prompt TuningCheng Han, Qifan Wang, Yiming Cui, Zhiwen Cao 等ICCV 2023 · 被引用 108 次
- ArtiBoost: Boosting Articulated 3D Hand-Object Pose Estimation via Online Exploration and SynthesisLixin Yang, Kailin Li, Xinyu Zhan, Jun Lv 等CVPR 2022 · 被引用 82 次
- DoubleField: Bridging the Neural Surface and Radiance Fields for High-fidelity Human Reconstruction and RenderingRuizhi Shao, Hongwen Zhang, He Zhang, Mingjia Chen 等CVPR 2022 · 被引用 72 次
它引用的顶会 Paper3
- Exploiting Spatial-Temporal Relationships for 3D Pose Estimation via Graph Convolutional NetworksYujun Cai, Liuhao Ge, Jun Liu, Jianfei Cai 等ICCV 2019 · 被引用 504 次
- End-to-End Hand Mesh Recovery From a Monocular RGB ImageXiong Zhang, Qiang Li, Hong Mo, Wenbo Zhang 等ICCV 2019 · 被引用 248 次
- HOPE-Net: A Graph-Based Model for Hand-Object Pose EstimationBardia Doosti, Shujon Naha, Majid Mirbagheri, David J. CrandallCVPR 2020
相关 Paper
- Learning Context with Priors for 3D Interacting Hand-Object Pose EstimationZengsheng Kuang, Changxing Ding, Huan YaoACM MM 2024 · 被引用 1 次
- A2J-Transformer: Anchor-to-Joint Transformer Network for 3D Interacting Hand Pose Estimation from a Single RGB ImageChanglong Jiang, Yang Xiao, Cunlin Wu, Mingyang Zhang 等CVPR 2023
- HandOccNet: Occlusion-Robust 3D Hand Mesh Estimation NetworkJoonKyu Park, Yeonguk Oh, Gyeongsik Moon, Hongsuk Choi 等CVPR 2022 · 被引用 116 次
- Keypoint Transformer: Solving Joint Identification in Challenging Hands and Object Interactions for Accurate 3D Pose EstimationShreyas Hampali, Sayan Deb Sarkar, Mahdi Rad, Vincent LepetitCVPR 2022 · 被引用 155 次
- Collaborative Learning for Hand and Object Reconstruction with Attention-guided Graph ConvolutionTze Ho Elden Tse, Kwang In Kim, Ales Leonardis, Hyung Jin ChangCVPR 2022 · 被引用 56 次
