Goldilocks Test Sets for Face Verification
Haiyu Wu, Sicong Tian, Aman Bhatta, Jacob Gutierrez, Grace Bezold, Genesis Argueta, Karl Ricanek, Michael C. King, Kevin W. Bowyer
Abstract
Reported face verification accuracy has reached a plateau on current well-known test sets. As a result, some difficult test sets have been assembled by reducing the image quality or adding artifacts to the image. However, we argue that test sets can be challenging without artificially reducing the image quality because the face recognition (FR) models suffer from correctly recognizing 1) the pairs from the same identity (i.e., genuine pairs) with a large face attribute difference, 2) the pairs from different identities (i.e., impostor pairs) with a small face attribute difference, and 3) the pairs of similar-looking identities (e.g., twins and relatives). We propose three challenging test sets to reveal important but ignored weaknesses of the existing FR algorithms. To challenge models on variation of facial attributes, we propose Hadrian and Eclipse to address facial hair differences and face exposure differences. The images in both test sets are high-quality and collected in a controlled environment. To challenge FR models on similar-looking persons, we propose ND-Twins, which contains images from a dedicated twins dataset. The LFW test protocol is used to structure the proposed test sets. Moreover, we introduce additional rules to assemble "Goldilocks 1 " level test sets, including 1) restricted number of occurrence of hard samples, 2) equal chance evaluation across demographic groups, and 3) constrained identity overlap across validation folds. Quantitatively, without further processing the images, the proposed test sets have on-par or higher difficulties than the existing test sets that add artifacts to the images. The datasets are available at: https://github.com/HaiyuWu/SOTA-Face- Recognition-Train-and-Test.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fa8b2abd-24bc-47d7-a877-95108f2b6f56Builds on8
- AdaFace: Quality Adaptive Margin for Face RecognitionMinchul Kim, Anil K. Jain, Xiaoming LiuCVPR 2022 · 509 citations
- SynFace: Face Recognition with Synthetic DataHaibo Qiu, Baosheng Yu, Dihong Gong, Zhifeng Li et al.ICCV 2021 · 162 citations
- UniFace: Unified Cross-Entropy Loss for Deep Face RecognitionJiancan Zhou, Xi Jia, Qiufu Li, Linlin Shen et al.ICCV 2023 · 38 citations
- DCFace: Synthetic Face Generation with Dual Condition Diffusion ModelMinchul Kim, Feng Liu, Anil K. Jain, Xiaoming LiuCVPR 2023
- MagFace: A Universal Representation for Face Recognition and Quality AssessmentQiang Meng, Shichao Zhao, Zhida Huang, Feng ZhouCVPR 2021
Related papers
- AHAN: Asymmetric Hierarchical Attention Network for Identical Twin Face VerificationHoang-Nhat NguyenAAAI 2026
- Benchmarking Algorithmic Bias in Face Recognition: An Experimental Approach Using Synthetic Faces and Human EvaluationHao Liang, Pietro Perona, Guha BalakrishnanICCV 2023 · 33 citations
- EFHQ: Multi-Purpose ExtremePose-Face-HQ DatasetTrung Tuan Dao, Duc Hong Vu, Cuong Pham, Anh Tuan TranCVPR 2024
- Vec2Face: Scaling Face Dataset Generation with Loosely Constrained VectorsHaiyu Wu, Jaskirat Singh, Sicong Tian, Liang Zheng et al.ICLR 2025
- HyperFace: Generating Synthetic Face Recognition Datasets by Exploring Face Embedding HypersphereHatef Otroshi-Shahreza, Sébastien MarcelICLR 2025
