HOT-Net: Non-Autoregressive Transformer for 3D Hand-Object Pose Estimation
Lin Huang, Jianchao Tan, Jingjing Meng, Ji Liu, Junsong Yuan
Abstract
As we use our hands frequently in daily activities, the analysis of hand-object interactions plays a critical role to many multimedia understanding and interaction applications. Different from conventional 3D hand-only and object-only pose estimation, estimating 3D hand-object pose is more challenging due to the mutual occlusions between hand and object, as well as the physical constraints between them. To overcome these issues, we propose to fully utilize the structural correlations among hand joints and object corners in order to obtain more reliable poses. Our work is inspired by structured output learning models in sequence transduction field like Transformer encoder-decoder framework. Besides modeling inherent dependencies from extracted 2D hand-object pose, our proposed Hand-Object Transformer Network (HOT-Net) also captures the structural correlations among 3D hand joints and object corners. Similar to Transformer's autoregressive decoder, by considering structured output patterns, this helps better constrain the output space and leads to more robust pose estimation. However, different from Transformer's sequential modeling mechanism, HOT-Net adopts a novel non-autoregressive decoding strategy for 3D hand-object pose estimation. Specifically, our model removes the Transformer's dependence on previously generated results and explicitly feeds a reference 3D hand-object pose into the decoding process to provide equivalent target pose patterns for parallely localizing each 3D keypoint. To further improve physical validity of estimated hand pose, besides anatomical constraints, we propose a cooperative pose constraint, aiming to enable the hand pose to cooperate with hand shape, to generate hand mesh. We demonstrate real-time speed and state-of-the-art performance on benchmark hand-object datasets for both 3D hand and object poses.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e4e770e-3a7b-4427-8321-cee4e29f857eCited by top-tier papers15
- Not All Tokens Are Equal: Human-centric Visual Analysis via Token Clustering TransformerWang Zeng, Sheng Jin, Wentao Liu, Chen Qian et al.CVPR 2022 · 132 citations
- DeepMultiCap: Performance Capture of Multiple Characters Using Sparse Multiview CamerasYang Zheng, Ruizhi Shao, Yuxiang Zhang, Tao Yu et al.ICCV 2021 · 112 citations
- E2VPT: An Effective and Efficient Approach for Visual Prompt TuningCheng Han, Qifan Wang, Yiming Cui, Zhiwen Cao et al.ICCV 2023 · 108 citations
- ArtiBoost: Boosting Articulated 3D Hand-Object Pose Estimation via Online Exploration and SynthesisLixin Yang, Kailin Li, Xinyu Zhan, Jun Lv et al.CVPR 2022 · 82 citations
- DoubleField: Bridging the Neural Surface and Radiance Fields for High-fidelity Human Reconstruction and RenderingRuizhi Shao, Hongwen Zhang, He Zhang, Mingjia Chen et al.CVPR 2022 · 72 citations
Builds on3
- Exploiting Spatial-Temporal Relationships for 3D Pose Estimation via Graph Convolutional NetworksYujun Cai, Liuhao Ge, Jun Liu, Jianfei Cai et al.ICCV 2019 · 504 citations
- End-to-End Hand Mesh Recovery From a Monocular RGB ImageXiong Zhang, Qiang Li, Hong Mo, Wenbo Zhang et al.ICCV 2019 · 248 citations
- HOPE-Net: A Graph-Based Model for Hand-Object Pose EstimationBardia Doosti, Shujon Naha, Majid Mirbagheri, David J. CrandallCVPR 2020
Related papers
- Learning Context with Priors for 3D Interacting Hand-Object Pose EstimationZengsheng Kuang, Changxing Ding, Huan YaoACM MM 2024 · 1 citation
- A2J-Transformer: Anchor-to-Joint Transformer Network for 3D Interacting Hand Pose Estimation from a Single RGB ImageChanglong Jiang, Yang Xiao, Cunlin Wu, Mingyang Zhang et al.CVPR 2023
- HandOccNet: Occlusion-Robust 3D Hand Mesh Estimation NetworkJoonKyu Park, Yeonguk Oh, Gyeongsik Moon, Hongsuk Choi et al.CVPR 2022 · 116 citations
- Keypoint Transformer: Solving Joint Identification in Challenging Hands and Object Interactions for Accurate 3D Pose EstimationShreyas Hampali, Sayan Deb Sarkar, Mahdi Rad, Vincent LepetitCVPR 2022 · 155 citations
- Collaborative Learning for Hand and Object Reconstruction with Attention-guided Graph ConvolutionTze Ho Elden Tse, Kwang In Kim, Ales Leonardis, Hyung Jin ChangCVPR 2022 · 56 citations
