FTAFace: Context-enhanced Face Detector with Fine-grained Task Attention
Deyu Wang, Dongchao Wen, Wei Tao, Lingxiao Yin, Tse-Wei Chen, Tadayuki Ito, Kinya Osa, Masami Kato
Abstract
In face detection, it is a common strategy to treat samples differently according to their difficulty for balancing training data distribution. However, we observe that widely used sampling strategies, such as OHEM and Focal loss, can lead to the performance imbalance between different tasks (e.g., classification and localization). Through analysis, we point out that, due to the driving of classification information, these sample-based strategies are difficult to coordinate the attention of different tasks during the training, thus leading to the above imbalance. Accordingly, we first confirm this by shifting the attention from the sample level to the task level. Then, we propose a fine-grained task attention method, a.k.a FTA, including inter-task importance and intra-task importance, which adaptively adjusts the attention of each item in the task from both global and local perspectives, so as to achieve finer optimization. In addition, we introduce transformer as a feature enhancer to assist our convolution network, and propose a context enhancement transformer, a.k.a CET, to mine the spatial relationship in the features towards more robust feature representation. Extensive experiments on WiderFace and FDDB benchmarks demonstrate that our method significantly boosts the baseline performance by 2.7%, 2.3% and 4.9% on easy, medium and hard validation sets respectively. Furthermore, the proposed FTAFace-light achieves higher accuracy than the state-of-the-art and reduces the amount of computation by 28.9%.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Focal Attention for Long-Range Interactions in Vision TransformersJianwei Yang, Chunyuan Li, Pengchuan Zhang, Xiyang Dai et al.NeurIPS 2021 · 228 citations
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski et al.AAAI 2022 · 529 citations
- PossLoss: A Reliable and Sensitive Facial Landmark Detection Loss FunctionQikui ZhuICCV 2025 · 1 citation
- Task Decoupled Knowledge Distillation For Lightweight Face DetectorsXiaoqing Liang, Xu Zhao, Chaoyang Zhao, Nanfei Jiang et al.ACM MM 2020 · 5 citations
- Balanced Hierarchical Contrastive Learning with Decoupled Queries for Fine-grained Object Detection in Remote Sensing ImagesJingzhou Chen, Dexin Chen, Fengchao Xiong, Yuntao Qian et al.CVPR 2026
