TokenHPE: Learning Orientation Tokens for Efficient Head Pose Estimation via Transformers
Cheng Zhang, Hai Liu, Yongjian Deng, Bochen Xie, Youfu Li
Abstract
Head pose estimation (HPE) has been widely used in the fields of human machine interaction, self-driving, and attention estimation. However, existing methods cannot deal with extreme head pose randomness and serious occlusions. To address these challenges, we identify three cues from head images, namely, neighborhood similarities, significant facial changes, and critical minority relationships. To leverage the observed findings, we propose a novel critical minority relationship-aware method based on the Transformer architecture in which the facial part relationships can be learned. Specifically, we design several orientation tokens to explicitly encode the basic orientation regions. Meanwhile, a novel token guide multiloss function is designed to guide the orientation tokens as they learn the desired regional similarities and relationships. We evaluate the proposed method on three challenging benchmark HPE datasets. Experiments show that our method achieves better performance compared with state-of-the-art methods. Our code is publicly available at https://github.com/zc2023/TokenHPE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 55b60b9c-dda9-4cc1-bc7d-e2f71a95ebf9Cited by top-tier papers2
- BAH Dataset for Ambivalence/Hesitancy Recognition in Videos for Digital Behavioural ChangeManuela González-González, Soufiane Belharbi, Muhammad Osama Zeeshan, Masoumeh Sharafi et al.ICLR 2026 · 18 citations
- FaceXFormer: A Unified Transformer for Facial AnalysisKartik Narayan, Vibashan VS, Rama Chellappa, Vishal M. PatelICCV 2025 · 16 citations
Builds on11
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- On the Relationship between Self-Attention and Convolutional LayersJean-Baptiste Cordonnier, Andreas Loukas, Martin JaggiICLR 2020 · 629 citations
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski et al.AAAI 2022 · 529 citations
Related papers
- Pose-Oriented Transformer with Uncertainty-Guided Refinement for 2D-to-3D Human Pose EstimationHan Li, Bowen Shi, Wenrui Dai, Hongwei Zheng et al.AAAI 2023 · 76 citations
- Direction Relation Transformer for Image CaptioningZeliang Song, Xiaofei Zhou, Linhua Dong, Jianlong Tan et al.ACM MM 2021 · 31 citations
- Densemarks: Learning Canonical Embeddings for Human Heads Images via Point TracksDmitrii Pozdeev, Alexey Artemov, Ananta R. Bhattarai, Artem SevastopolskyICLR 2026
- Towards Accurate Facial Landmark Detection via Cascaded TransformersHui Li, Zidong Guo, Seon-Min Rhee, Seungju Han et al.CVPR 2022 · 45 citations
- Pose-guided Inter- and Intra-part Relational Transformer for Occluded Person Re-IdentificationZhongxing Ma, Yifan Zhao, Jia LiACM MM 2021 · 66 citations
