Self-Supervised Emotion Representation Disentanglement for Speech-Preserving Facial Expression Manipulation
Zhihua Xu, Tianshui Chen, Zhijing Yang, Chunmei Qing, Yukai Shi, Liang Lin
Abstract
Speech-preserving Facial Expression Manipulation (SPFEM) aims to alter facial emotions in video content while preserving the facial movements associated with speech. Current works often fall short due to the inadequate representation of emotion as well as the absence of time-aligned paired data-two corresponding frames from the same speaker that showcase the same speech content but differ in emotional expression. In this work, we introduce a novel framework, Self-Supervised Emotion Representation Disentanglement (SSERD), to disentangle emotion representation for accurate emotion transfer while implementing a paired data construction module to facilitate automated, photorealistic facial animations. Specifically, We developed a module for learning emotion latent codes using StyleGAN's latent space, employing a cross-attention mechanism to extract and predict emotion editing codes, with contrastive learning to differentiate emotions. To overcome the lack of strictly paired data in the SPFEM task, we exploit pretrained StyleGAN to generate paired data, focusing on expression vectors unrelated to mouth shape. Additionally, we employed a hybrid training strategy using both synthetic paired and real unpaired data to enhance the realism of SPFEM model's generated images. Extensive experiments conducted on benchmark datasets, including MEAD and RAVDESS, have validated the effectiveness of our framework, demonstrating its superior capability in generating photorealistic and expressive facial animations.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c6d1288d-9470-4ef8-97d9-6e6fc9449be5Cited by top-tier papers1
Ask how each one uses itRelated papers
- Unsupervised Disentanglement of Linear-Encoded Facial SemanticsYutong Zheng, Yu-Kai Huang, Ran Tao, Zhiqiang Shen et al.CVPR 2021
- SegTalker: Segmentation-based Talking Face Generation with Mask-guided Local EditingLingyu Xiong, Xize Cheng, Jintao Tan, Xianjia Wu et al.ACM MM 2024 · 10 citations
- Learning Adaptive Spatial Coherent Correlations for Speech-Preserving Facial Expression ManipulationTianshui Chen, Jianman Lin, Zhijing Yang, Chunmei Qing et al.CVPR 2024
- A Latent Transformer for Disentangled Face Editing in Images and VideosXu Yao, Alasdair Newson, Yann Gousseau, Pierre HellierICCV 2021 · 97 citations
- SD-GAN: Semantic Decomposition for Face Image Synthesis with Discrete AttributeKangneng Zhou, Xiaobin Zhu, Daiheng Gao, Kai Lee et al.ACM MM 2022 · 2 citations
