A2J-Transformer: Anchor-to-Joint Transformer Network for 3D Interacting Hand Pose Estimation from a Single RGB Image
Changlong Jiang, Yang Xiao, Cunlin Wu, Mingyang Zhang, Jinghong Zheng, Zhiguo Cao, Joey Tianyi Zhou
摘要
3D interacting hand pose estimation from a single RGB image is a challenging task, due to serious self-occlusion and inter-occlusion towards hands, confusing similar appearance patterns between 2 hands, ill-posed joint position mapping from 2D to 3D, etc.. To address these, we propose to extend A2J-the state-of-the-art depth-based 3D single hand pose estimation method-to RGB domain under interacting hand condition. Our key idea is to equip A2J with strong local-global aware ability to well capture interacting hands' local fine details and global articulated clues among joints jointly. To this end, A2J is evolved under Transformer's non-local encoding-decoding framework to build A2J-Transformer. It holds 3 main advantages over A2J. First, self-attention across local anchor points is built to make them global spatial context aware to better capture joints' articulation clues for resisting occlusion. Secondly, each anchor point is regarded as learnable query with adaptive feature learning for facilitating pattern fitting capacity, instead of having the same local representation with the others. Last but not least, anchor point locates in 3D space instead of 2D as in A2J, to leverage 3D pose prediction. Experiments on challenging InterHand 2.6M demonstrate that, A2J-Transformer can achieve state-ofthe-art model-free performance (3.38mm MPJPE advancement in 2-hand case) and can also be applied to depth domain with strong generalization. The code is avaliable at https://github.com/ChanglongJiangGit/ A2J-Transformer .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- TACO: Benchmarking Generalizable Bimanual Tool-ACtion-Object UnderstandingYun Liu, Haolin Yang, Xu Si, Ling Liu 等CVPR 2024 · 被引用 12 次
- Initializing Variable-sized Vision Transformers from Learngene with Learnable TransformationShiyu Xia, Yuankun Zu, Xu Yang, Xin GengNeurIPS 2024 · 被引用 9 次
- Touchscreen-based Hand Tracking for Remote Whiteboard InteractionXinshuang Liu, Yizhong Zhang, Xin TongUIST 2024 · 被引用 8 次
- QORT-Former: Query-optimized Real-time Transformer for Understanding Two Hands Manipulating ObjectsElkhan Ismayilzada, MD Khalequzzaman Chowdhury Sayem, Yihalem Yimolal Tiruneh, Mubarrat Tajoar Chowdhury 等AAAI 2025 · 被引用 4 次
- PAD-Hand: Physics-Aware Diffusion for Hand Motion RecoveryElkhan Ismayilzada, Yufei Zhang, Zijun CuiCVPR 2026 · 被引用 4 次
它引用的顶会 Paper12
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang 等ICLR 2022 · 被引用 1,218 次
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng 等ICCV 2021 · 被引用 974 次
- Mesh GraphormerKevin Lin, Lijuan Wang, Zicheng LiuICCV 2021 · 被引用 399 次
- A2J: Anchor-to-Joint Regression Network for 3D Articulated Pose Estimation From a Single Depth ImageFu Xiong, Boshen Zhang, Yang Xiao, Zhiguo Cao 等ICCV 2019 · 被引用 178 次
相关 Paper
- Keypoint Transformer: Solving Joint Identification in Challenging Hands and Object Interactions for Accurate 3D Pose EstimationShreyas Hampali, Sayan Deb Sarkar, Mahdi Rad, Vincent LepetitCVPR 2022 · 被引用 155 次
- HOT-Net: Non-Autoregressive Transformer for 3D Hand-Object Pose EstimationLin Huang, Jianchao Tan, Jingjing Meng, Ji Liu 等ACM MM 2020 · 被引用 59 次
- Interacting Two-Hand 3D Pose and Shape Reconstruction from Single Color ImageBaowen Zhang, Yangang Wang, Xiaoming Deng, Yinda Zhang 等ICCV 2021 · 被引用 114 次
- Decoupled Iterative Refinement Framework for Interacting Hands Reconstruction from a Single RGB ImagePengfei Ren, Chao Wen, Xiaozheng Zheng, Zhou Xue 等ICCV 2023 · 被引用 15 次
- Multiple View Geometry Transformers for 3D Human Pose EstimationZiwei Liao, Jialiang Zhu, Chunyu Wang, Han Hu 等CVPR 2024
