FaceInsight: A Multimodal Large Language Model for Face Perception
Jingzhi Li, Changjiang Luo, Ruoyu Chen, Hua Zhang, Wenqi Ren, Jianhou Gan, Xiaochun Cao
摘要
Can you give a detailed description of the person in the image ? Question FaceInsight: The individual is a white male, aged 50-59. He has a square face, a fully visible forehead, and a double chin. His hair is gray with a receding hairline, and he has a happy facial expression. His eyebrows are bushy. His eyes are brown with noticeable bags underneath and no eyewear. His nose is big and pointy. His mouth is closed. He has no beard. He is wearing a necktie. Qwen2.5-VL-7B-Instruct: The individual is a white male, aged 40-49. He has an oval face, high cheekbones, rosy cheeks, pale skin, and a chubby appearance with a double chin. His hair is dark, brown, and gray, with a receding hairline and sideburns. His facial expression is neutral. His eyebrows are arched and bushy. His eyes are narrow with bags underneath and no eyewear. His nose is big and pointy. His mouth is slightly open and has big lips. He has a 5 o'clock shadow and a mustache but no beard. He is wearing lipstick and a necktie. He is described as attractive and smiling. MiniCPM-Llama3-v2.5: The individual is a white male, aged 50-59. He has an oval face with a double chin and gray hair. He has sideburns, a mustache, and a 5 o'clock shadow but no beard. His facial expression is neutral. He has bags under eyes.
The individual is a white male, aged 50-59. He has gray hair and a happy facial expression.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation ModelsZhicheng Zhang, Weicheng Wang, Yongjie Zhu, Wenyu Qin 等NeurIPS 2025 · 被引用 11 次
- UniFace: A fied ine-grained Understanding and Generation ModelJunzhe Li, Sifan Zhou, Liya Guo, Xuerui Qiu 等ICLR 2026
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
相关 Paper
- CelebV-Text: A Large-Scale Facial Text-Video DatasetJianhui Yu, Hao Zhu, Liming Jiang, Chen Change Loy 等CVPR 2023
- NeckFace: Continuously Tracking Full Facial Expressions on Neck-mounted WearablesTuochao Chen, Yaxuan Li, Songyun Tao, Hyunchul Lim 等UbiComp 2021 · 被引用 23 次
- FaceBench: A Multi-View Multi-Level Facial Attribute VQA Dataset for Benchmarking Face Perception MLLMsXiaoqin Wang, Xusen Ma, Xianxu Hou, Meidan Ding 等CVPR 2025
- Collaborative Diffusion for Multi-Modal Face Generation and EditingZiqi Huang, Kelvin C. K. Chan, Yuming Jiang, Ziwei LiuCVPR 2023
- MMHead: Towards Fine-grained Multi-modal 3D Facial AnimationSijing Wu, Yunhao Li, Yichao Yan, Huiyu Duan 等ACM MM 2024 · 被引用 17 次
