FaceInsight: A Multimodal Large Language Model for Face Perception
Jingzhi Li, Changjiang Luo, Ruoyu Chen, Hua Zhang, Wenqi Ren, Jianhou Gan, Xiaochun Cao
Abstract
Can you give a detailed description of the person in the image ? Question FaceInsight: The individual is a white male, aged 50-59. He has a square face, a fully visible forehead, and a double chin. His hair is gray with a receding hairline, and he has a happy facial expression. His eyebrows are bushy. His eyes are brown with noticeable bags underneath and no eyewear. His nose is big and pointy. His mouth is closed. He has no beard. He is wearing a necktie. Qwen2.5-VL-7B-Instruct: The individual is a white male, aged 40-49. He has an oval face, high cheekbones, rosy cheeks, pale skin, and a chubby appearance with a double chin. His hair is dark, brown, and gray, with a receding hairline and sideburns. His facial expression is neutral. His eyebrows are arched and bushy. His eyes are narrow with bags underneath and no eyewear. His nose is big and pointy. His mouth is slightly open and has big lips. He has a 5 o'clock shadow and a mustache but no beard. He is wearing lipstick and a necktie. He is described as attractive and smiling. MiniCPM-Llama3-v2.5: The individual is a white male, aged 50-59. He has an oval face with a double chin and gray hair. He has sideburns, a mustache, and a 5 o'clock shadow but no beard. His facial expression is neutral. He has bags under eyes.
The individual is a white male, aged 50-59. He has gray hair and a happy facial expression.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation ModelsZhicheng Zhang, Weicheng Wang, Yongjie Zhu, Wenyu Qin et al.NeurIPS 2025 · 11 citations
- UniFace: A fied ine-grained Understanding and Generation ModelJunzhe Li, Sifan Zhou, Liya Guo, Xuerui Qiu et al.ICLR 2026
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- CelebV-Text: A Large-Scale Facial Text-Video DatasetJianhui Yu, Hao Zhu, Liming Jiang, Chen Change Loy et al.CVPR 2023
- NeckFace: Continuously Tracking Full Facial Expressions on Neck-mounted WearablesTuochao Chen, Yaxuan Li, Songyun Tao, Hyunchul Lim et al.UbiComp 2021 · 23 citations
- FaceBench: A Multi-View Multi-Level Facial Attribute VQA Dataset for Benchmarking Face Perception MLLMsXiaoqin Wang, Xusen Ma, Xianxu Hou, Meidan Ding et al.CVPR 2025
- Collaborative Diffusion for Multi-Modal Face Generation and EditingZiqi Huang, Kelvin C. K. Chan, Yuming Jiang, Ziwei LiuCVPR 2023
- MMHead: Towards Fine-grained Multi-modal 3D Facial AnimationSijing Wu, Yunhao Li, Yichao Yan, Huiyu Duan et al.ACM MM 2024 · 17 citations
