Vec2Face: Scaling Face Dataset Generation with Loosely Constrained Vectors
Haiyu Wu, Jaskirat Singh, Sicong Tian, Liang Zheng, Kevin W. Bowyer
Abstract
This paper studies how to synthesize face images of non-existent persons, to create a dataset that allows effective training of face recognition (FR) models. Besides generating realistic face images, two other important goals are: 1) the ability to generate a large number of distinct identities (inter-class separation), and 2) a proper variation in appearance of the images for each identity (intra-class variation). However, existing works 1) are typically limited in how many well-separated identities can be generated and 2) either neglect or use an external model for attribute augmentation. We propose Vec2Face, a holistic model that uses only a sampled vector as input and can flexibly generate and control the identity of face images and their attributes. Composed of a feature masked autoencoder and an image decoder, Vec2Face is supervised by face image reconstruction and can be conveniently used in inference. Using vectors with low similarity among themselves as inputs, Vec2Face generates well-separated identities. Randomly perturbing an input identity vector within a small range allows Vec2Face to generate faces of the same identity with proper variation in face attributes. It is also possible to generate images with designated attributes by adjusting vector values with a gradient descent method. Vec2Face has efficiently synthesized as many as 300K identities, whereas 60K is the largest number of identities created in the previous works. As for performance, FR models trained with the generated HSFace datasets, from 10k to 300k identities, achieve state-of-the-art accuracy, from 92% to 93.52%, on five real-world test sets (i.e., LFW, CFP-FP, AgeDB-30, CALFW, and CPLFW). For the first time, the FR model trained using our synthetic training set achieves higher accuracy than that trained using a same-scale training set of real face images on the CALFW, IJBB, and IJBC test sets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d1f4dc9f-3d8a-4898-b288-b4fa628237e1Cited by top-tier papers7
- MR-FIQA: Face Image Quality Assessment with Multi-Reference Representations from Synthetic Data GenerationFu-Zhao Ou, Chongyi Li, Shiqi Wang, Sam KwongICCV 2025 · 4 citations
- LSAP-PV: High-Fidelity Palm Vein Image Synthesis via Layered Spectral Absorption Projection-Guided Diffusion ModelSheng Shang, Chenglong Zhao, Ruixin Zhang, Jianlong Jin et al.AAAI 2026 · 1 citation
- VIGFace: Virtual Identity Generation for Privacy-Free Face Recognition DatasetMinsoo Kim, Min-Cheol Sagong, Gi Pyo Nam, Junghyun Cho et al.ICCV 2025 · 1 citation
- Diff-Palm: Realistic Palmprint Generation with Polynomial Creases and Intra-Class Variation Controllable Diffusion ModelsJianlong Jin, Chenglong Zhao, Ruixin Zhang, Sheng Shang et al.CVPR 2025
- From Measurement to Mitigation: Quantifying and Reducing Identity Leakage in Image Representation Encoders with Linear Subspace RemovalDaniel George, Charles Yeh, Daniel Lee, Yifei ZhangCVPR 2026
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- SynFace: Face Recognition with Synthetic DataHaibo Qiu, Baosheng Yu, Dihong Gong, Zhifeng Li et al.ICCV 2021 · 162 citations
- BlendFace: Re-designing Identity Encoders for Face-SwappingKaede Shiohara, Xingchao Yang, Takafumi TaketomiICCV 2023 · 83 citations
- IDiff-Face: Synthetic-based Face Recognition through Fizzy Identity-Conditioned Diffusion ModelsFadi Boutros, Jonas Henry Grebe, Arjan Kuijper, Naser DamerICCV 2023 · 106 citations
- How to Boost Face Recognition with StyleGAN?Artem Sevastopolsky, Yury Malkov, Nikita Durasov, Luisa Verdoliva et al.ICCV 2023 · 17 citations
- UIFace: Unleashing Inherent Model Capabilities to Enhance Intra-Class Diversity in Synthetic Face RecognitionXiao Lin, Yuge Huang, Jianqing Xu, Yuxi Mi et al.ICLR 2025
