Sparse Local Patch Transformer for Robust Face Alignment and Landmarks Inherent Relation Learning
Jiahao Xia, Weiwei Qu, Wenjian Huang, Jianguo Zhang, Xi Wang, Min Xu
Abstract
Heatmap regression methods have dominated face alignment area in recent years while they ignore the inherent relation between different landmarks. In this paper, we propose a Sparse Local Patch Transformer (SLPT) for learning the inherent relation. The SLPT generates the representation of each single landmark from a local patch and aggregates them by an adaptive inherent relation based on the attention mechanism. The subpixel coordinate of each landmark is predicted independently based on the aggregated feature. Moreover, a coarse-to-fine framework is further introduced to incorporate with the SLPT, which enables the initial landmarks to gradually converge to the target facial landmarks using fine-grained features from dynamically resized local patches. Extensive experiments carried out on three popular benchmarks, including WFLW, 300W and COFW, demonstrate that the proposed method works at the state-of-the-art level with much less computational complexity by learning the inherent relation between facial landmarks. The code is available at the project website <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> https://github.com/Jiahao-UTS/SLPT-master.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 55d67bf2-3aff-40cf-b0cc-c3ce09c380f9Cited by top-tier papers13
- Learning Motion-Robust Remote Photoplethysmography through Arbitrary Resolution VideosJianwei Li, Zitong Yu, Jingang ShiAAAI 2023 · 65 citations
- FaceXFormer: A Unified Transformer for Facial AnalysisKartik Narayan, Vibashan VS, Rama Chellappa, Vishal M. PatelICCV 2025 · 16 citations
- KeyPosS: Plug-and-Play Facial Landmark Detection through GPS-Inspired True-Range MultilaterationXu Bao, Zhi-Qi Cheng, Jun-Yan He, Wangmeng Xiang et al.ACM MM 2023 · 5 citations
- FacialFlowNet: Advancing Facial Optical Flow Estimation with a Diverse Dataset and a Decomposed ModelJianzhi Lu, Ruian He, Shili Zhou, Weimin Tan et al.ACM MM 2024 · 4 citations
- POPoS: Improving Efficient and Robust Facial Landmark Detection with Parallel Optimal Position SearchChong-Yang Xiang, Jun-Yan He, Zhi-Qi Cheng, Xiao Wu et al.AAAI 2025 · 3 citations
Builds on11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Adaptive Wing Loss for Robust Face Alignment via Heatmap RegressionXinyao Wang, Liefeng Bo, Fuxin LiICCV 2019 · 293 citations
- DeCaFA: Deep Convolutional Cascade for Face Alignment in the WildArnaud Dapogny, Matthieu Cord, Kevin BaillyICCV 2019 · 91 citations
- Aggregation via Separation: Boosting Facial Landmark Detector With Semi-Supervised Style TranslationShengju Qian, Keqiang Sun, Wayne Wu, Chen Qian et al.ICCV 2019 · 79 citations
Related papers
- Towards Accurate Facial Landmark Detection via Cascaded TransformersHui Li, Zidong Guo, Seon-Min Rhee, Seungju Han et al.CVPR 2022 · 45 citations
- Attentive One-Dimensional Heatmap Regression for Facial Landmark Detection and TrackingShi Yin, Shangfei Wang, Xiaoping Chen, Enhong Chen et al.ACM MM 2020 · 22 citations
- FreeEnricher: Enriching Face Landmarks without Additional CostYangyu Huang, Xi Chen, Jongyoo Kim, Hao Yang et al.AAAI 2023 · 3 citations
- Knowing When to Quit: Selective Cascaded Regression with Patch Attention for Real-Time Face AlignmentGil Shapira, Noga Levy, Ishay Goldin, Roy Josef JevnisekACM MM 2021 · 3 citations
- Dual Focus-Attention Transformer for Robust Point Cloud RegistrationKexue Fu, Mingzhi Yuan, Changwei Wang, Weiguang Pang et al.CVPR 2025
