FLIP-80M: 80 Million Visual-Linguistic Pairs for Facial Language-Image Pre-Training
Yudong Li, Xianxu Hou, Dezhi Zheng, Linlin Shen, Zhe Zhao
Abstract
While significant progress has been made in multi-modal learning driven by large-scale image-text datasets, there is still a noticeable gap in the availability of such datasets within the facial domain. To facilitate and advance the field of facial representation learning, we present FLIP-80M, a large-scale visual-linguistic dataset comprising over 80 million face images paired with text descriptions. FLIP-80M is constructed by leveraging the large openly available image-text-pair dataset LAION-5B and a mixed-method approach to filter face-related pairs from both visual and linguistic perspectives. Our curation process involves face detection, face caption classification, text de-noising, and synthesis-based image augmentation. As a result, FLIP-80M stands as the largest face-text dataset to date. To evaluate the potential of our dataset, we fine-tune the CLIP model using the proposed FLIP-80M, to create FLIP (Facial Language-Image Pretraining) and assess its representation capabilities across various downstream tasks. Our experiments demonstrate that our FLIP model achieves state-of-the-art results in a range of face analysis tasks, including face parsing, face alignment, and face attribute classification. The dataset and models are available at https://github.com/ydli-ai/FLIP.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 2c521680-d14c-42b5-8ba7-3b05b03ddd33Cited by top-tier papers6
- FineXtrol: Controllable Motion Generation via Fine-Grained TextKeming Shen, Bizhu Wu, Junliang Chen, Xiaoqin Wang et al.AAAI 2026 · 3 citations
- DisFaceRep: Representation Disentanglement for Co-occurring Facial Components in Weakly Supervised Face ParsingXiaoqin Wang, Xianxu Hou, Meidan Ding, Junliang Chen et al.ACM MM 2025 · 1 citation
- UniFace: A fied ine-grained Understanding and Generation ModelJunzhe Li, Sifan Zhou, Liya Guo, Xuerui Qiu et al.ICLR 2026
- CA-Edit: Causality-Aware Condition Adapter for High-Fidelity Local Facial Attribute EditingXiaole Xian, Xilin He, Zenghao Niu, Junliang Zhang et al.AAAI 2025
- Evading Data Provenance in Deep Neural NetworksHongyu Zhu, Sichu Liang, Wenwen Wang, Zhuomeng Zhang et al.ICCV 2025
Related papers
- Scaling Language-Image Pre-Training via MaskingYanghao Li, Haoqi Fan, Ronghang Hu, Christoph Feichtenhofer et al.CVPR 2023
- Vision Model Pre-training on Interleaved Image-Text Data via Latent Compression LearningChenyu Yang, Xizhou Zhu, Jinguo Zhu, Weijie Su et al.NeurIPS 2024 · 10 citations
- PixCLIP: Towards Fine-grained Vision-Language Understanding via Any-granularity Pixel-Text AlignmentYicheng Xiao, Yu Chen, Hao-Xuan Ma, Jiale Hong et al.ICML 2026 · 4 citations
- General Facial Representation Learning in a Visual-Linguistic MannerYinglin Zheng, Hao Yang, Ting Zhang, Jianmin Bao et al.CVPR 2022 · 161 citations
- Reproducible Scaling Laws for Contrastive Language-Image LearningMehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman et al.CVPR 2023
