F-Bench: Rethinking Human Preference Evaluation Metrics for Benchmarking Face Generation, Customization, and Restoration
Lu Liu, Huiyu Duan, Qiang Hu, Liu Yang, Chunlei Cai, Tianxiao Ye, Huayu Liu, Xiaoyun Zhang, Guangtao Zhai
Abstract
Artificial intelligence generative models exhibit remarkable capabilities in content creation, particularly in face image generation, customization, and restoration. However, current AI-generated faces (AIGFs) often fall short of human preferences due to unique distortions, unrealistic details, and unexpected identity shifts, underscoring the need for a comprehensive quality evaluation framework for AIGFs. To address this need, we introduce FaceQ, a large-scale, comprehensive database of AI-generated Face images with fine-grained Quality annotations reflecting human preferences. The FaceQ database comprises 12,255 images generated by 29 models across three tasks: (1) face generation, (2) face customization, and (3) face restoration. It includes 32,742 mean opinion scores (MOSs) from 180 annotators, assessed across multiple dimensions: quality, authenticity, identity (ID) fidelity, and text-image correspondence. Using the FaceQ database, we establish F-Bench, a benchmark for comparing and evaluating face generation, customization, and restoration models, highlighting strengths and weaknesses across various prompts and evaluation dimensions. Additionally, we assess the performance of existing image quality assessment (IQA), face quality assessment (FQA), AI-generated content image quality assessment (AIGCIQA), and preference evaluation metrics, manifesting that these standard metrics are relatively ineffective in evaluating authenticity, ID fidelity, and text-image correspondence. The FaceQ database will be publicly available upon publication.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Agentic Retoucher for Text-To-Image GenerationShaocheng Shen, Jianfeng Liang, Chunlei Cai, Cong Geng et al.CVPR 2026 · 9 citations
- Omni2: Unifying Omnidirectional Image Generation and Editing in an Omni ModelLiu Yang, Huiyu Duan, Yucheng Zhu, Xiaohong Liu et al.ACM MM 2025 · 2 citations
- Market-Bench: Benchmarking Large Language Models on Economic and Trade CompetitionYushuo Zheng, Huiyu Duan, Zicheng Zhang, Yucheng Zhu et al.ACL 2026 · 1 citation
Builds on40
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- Multi-Dimensional Text-to-Face Image Quality Assessment Using LLM: Database and MethodYixuan Gao, Xiongkuo Min, Jinliang Han, Yuqin Cao et al.ACM MM 2025 · 2 citations
- Fine-grained Image Quality Assessment for Perceptual Image RestorationXiangfei Sheng, Xiaofeng Pan, Zhichao Yang, Pengfei Chen et al.AAAI 2026 · 4 citations
- Rethinking Deep Face RestorationYang Zhao, Yu-Chuan Su, Chun-Te Chu, Yandong Li et al.CVPR 2022 · 20 citations
- LMME3DHF: Benchmarking and Evaluating Multimodal 3D Human Face Generation with LMMsWoo Yi Yang, Jiarui Wang, Sijing Wu, Huiyu Duan et al.ACM MM 2025 · 7 citations
- MR-FIQA: Face Image Quality Assessment with Multi-Reference Representations from Synthetic Data GenerationFu-Zhao Ou, Chongyi Li, Shiqi Wang, Sam KwongICCV 2025 · 4 citations
