Talk2Face: A Unified Sequence-based Framework for Diverse Face Generation and Analysis Tasks
Yudong Li, Xianxu Hou, Zhe Zhao, Linlin Shen, Xuefeng Yang, Kimmo Yan
摘要
Facial analysis is an important domain in computer vision and has received extensive research attention. For numerous downstream tasks with different input/output formats and modalities, existing methods usually design task-specific architectures and train them using face datasets collected in the particular task domain. In this work, we proposed a single model, Talk2Face, to simultaneously tackle a large number of face generation and analysis tasks, e.g. text guided face synthesis, face captioning and age estimation. Specifically, we cast different tasks into a sequence-to-sequence format with the same architecture, parameters and objectives. While text and facial images are tokenized to sequences, the annotation labels of faces for different tasks are also converted to natural languages for unified representation. We collect a set of 2.3M face-text pairs from available datasets across different tasks, to train the proposed model. Uniform templates are then designed to enable the model to perform different downstream tasks, according to the task context and target. Experiments on different tasks show that our model achieves better face generation and caption performances than SOTA approaches. On age estimation and multi-attribute classification, our model reaches competitive performance with those models specially designed and trained for these particular tasks. In practice, our model is much easier to be deployed to different facial analysis related tasks. Code and dataset will be available at https://github.com/ydli-ai/Talk2Face.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- UMMAFormer: A Universal Multimodal-adaptive Transformer Framework for Temporal Forgery LocalizationRui Zhang, Hongxia Wang, Mingshan Du, Hanqing Liu 等ACM MM 2023 · 被引用 42 次
- Learning Profitable NFT Image Diffusions via Multiple Visual-Policy Guided Reinforcement LearningHuiguo He, Tianfu Wang, Huan Yang, Jianlong Fu 等ACM MM 2023 · 被引用 8 次
相关 Paper
- FaceXFormer: A Unified Transformer for Facial AnalysisKartik Narayan, Vibashan VS, Rama Chellappa, Vishal M. PatelICCV 2025 · 被引用 16 次
- FaceComposer: A Unified Model for Versatile Facial Content CreationJiayu Wang, Kang Zhao, Yifeng Ma, Shiwei Zhang 等NeurIPS 2023 · 被引用 14 次
- Bridging Facial Understanding and Animation via Language ModelsLuchuan Song, Pinxin Liu, Haiyang Liu, Zhenchao Jin 等CVPR 2026
- When Age-Invariant Face Recognition Meets Face Age Synthesis: A Multi-Task Learning FrameworkZhizhong Huang, Junping Zhang, Hongming ShanCVPR 2021
- Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual AwarenessJiaxing Zhao, Boyuan Sun, Xiang Chen, Xihan WeiAAAI 2026
