Attention-Driven Cropping for Very High Resolution Facial Landmark Detection
Prashanth Chandran, Derek Bradley, Markus Gross, Thabo Beeler
Abstract
Facial landmark detection is a fundamental task for many consumer and high-end applications and is almost entirely solved by machine learning methods today. Existing datasets used to train such algorithms are primarily made up of only low resolution images, and current algorithms are limited to inputs of comparable quality and resolution as the training dataset. On the other hand, high resolution imagery is becoming increasingly more common as consumer cameras improve in quality every year. Therefore, there is need for algorithms that can leverage the rich information available in high resolution imagery. Naïvely attempting to reuse existing network architectures on high resolution imagery is prohibitive due to memory bottlenecks on GPUs. The only current solution is to downsample the images, sacrificing resolution and quality. Building on top of recent progress in attention-based networks, we present a novel, fully convolutional regional architecture that is specially designed for predicting landmarks on very high resolution facial images without downsampling. We demonstrate the flexibility of our architecture by training the proposed model with images of resolutions ranging from 256 x 256 to 4K. In addition to being the first method for facial landmark detection on high resolution images, our approach achieves superior performance over traditional (holistic) state-of-the-art architectures across ALL resolutions, leading to a general-purpose, extremely flexible, high quality landmark detector.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1068cbd0-0a81-422b-9b2a-57f4d97c0635Cited by top-tier papers8
- Towards Accurate Facial Landmark Detection via Cascaded TransformersHui Li, Zidong Guo, Seon-Min Rhee, Seungju Han et al.CVPR 2022 · 45 citations
- Improving Robustness of Facial Landmark Detection by Defending against Adversarial AttacksCongcong Zhu, Xiaoqiang Li, Jide Li, Songmin DaiICCV 2021 · 34 citations
- Localization with Sampling-ArgmaxJiefeng Li, Tong Chen, Ruiqi Shi, Yujing Lou et al.NeurIPS 2021 · 25 citations
- KeyPosS: Plug-and-Play Facial Landmark Detection through GPS-Inspired True-Range MultilaterationXu Bao, Zhi-Qi Cheng, Jun-Yan He, Wangmeng Xiang et al.ACM MM 2023 · 5 citations
- POPoS: Improving Efficient and Robust Facial Landmark Detection with Parallel Optimal Position SearchChong-Yang Xiang, Jun-Yan He, Zhi-Qi Cheng, Xiao Wu et al.AAAI 2025 · 3 citations
Builds on2
Related papers
- Attentive One-Dimensional Heatmap Regression for Facial Landmark Detection and TrackingShi Yin, Shangfei Wang, Xiaoping Chen, Enhong Chen et al.ACM MM 2020 · 22 citations
- Joint Super-Resolution and Alignment of Tiny FacesYu Yin, Joseph P. Robinson, Yulun Zhang, Yun FuAAAI 2020 · 38 citations
- Heatmap Regression without Soft-Argmax for Facial Landmark DetectionChiao-An Yang, Raymond A. YehICCV 2025 · 3 citations
- Learning to Detect 3D Facial Landmarks via Heatmap Regression with Graph Convolutional NetworkYuan Wang, Min Cao, Zhenfeng Fan, Silong PengAAAI 2022 · 30 citations
- Towards High-Resolution Salient Object DetectionYi Zeng, Pingping Zhang, Zhe Lin, Jianming Zhang et al.ICCV 2019 · 232 citations
