Learning Adaptive Spatial Coherent Correlations for Speech-Preserving Facial Expression Manipulation
Tianshui Chen, Jianman Lin, Zhijing Yang, Chunmei Qing, Liang Lin
Abstract
Speech-preserving facial expression manipulation (SPFEM) aims to modify facial emotions while meticulously maintaining the mouth animation associated with spoken content. Current works depend on inaccessible paired training samples for the person, where two aligned frames exhibit the same speech content yet differ in emotional expression, limiting the SPFEM applications in real-world scenarios. In this work, we discover that speakers who convey the same content with different emotions exhibit highly correlated local facial animations, providing valuable supervision for SPFEM. To capitalize on this insight, we propose a novel adaptive spatial coherent correlation learning (ASCCL) algorithm, which models the aforementioned correlation as an explicit metric and integrates the metric to supervise manipulating facial expression and meanwhile better preserving the facial animation of spoken contents. To this end, it first learns a spatial coherent correlation metric, ensuring the visual disparities of adjacent local regions of the image belonging to one emotion are similar to those of the corresponding counterpart of the image belonging to another emotion. Recognizing that visual disparities are not uniform across all regions, we have also crafted a disparity-aware adaptive strategy that prioritizes regions that present greater challenges. During SPFEM model training, we construct the adaptive spatial coherent correlation metric between corresponding local regions of the input and output images as addition loss to supervise the generation * Zhijing Yang is the corresponding author. Tianshui Chen, Jianman Lin, and Zhijing Yang are with
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ef6ff51-8bb0-419a-8575-c416d7839a06Cited by top-tier papers1
Ask how each one uses itBuilds on10
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Alias-Free Generative Adversarial NetworksTero Karras, Miika Aittala, Samuli Laine, Erik Härkönen et al.NeurIPS 2021 · 2,126 citations
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- Learning an animatable detailed 3D face model from in-the-wild imagesYao Feng, Haiwen Feng, Michael J. Black, Timo BolkartSIGGRAPH 2021 · 662 citations
- Neural Emotion Director: Speech-preserving semantic control of facial expressions in "in-the-wild" videosFoivos Paraperas Papantoniou, Panagiotis Paraskevas Filntisis, Petros Maragos, Anastasios RoussosCVPR 2022 · 32 citations
Related papers
- Self-Supervised Emotion Representation Disentanglement for Speech-Preserving Facial Expression ManipulationZhihua Xu, Tianshui Chen, Zhijing Yang, Chunmei Qing et al.ACM MM 2024 · 5 citations
- JDMAN: Joint Discriminative and Mutual Adaptation Networks for Cross-Domain Facial Expression RecognitionYingjian Li, Yingnan Gao, Bingzhi Chen, Zheng Zhang et al.ACM MM 2021 · 21 citations
- Learning Motion Refinement for Unsupervised Face AnimationJiale Tao, Shuhang Gu, Wen Li, Lixin DuanNeurIPS 2023 · 10 citations
- BHGap: A Deep Iterative Prompting and Multi-stage Alignment Framework for Dynamic Facial Expression RecognitionYichi Zhang, Yunqi Han, Jiayue Ding, Liangyu ChenWWW 2026
- Contrastive Adversarial Learning for Person Independent Facial Emotion RecognitionDae Ha Kim, Byung Cheol SongAAAI 2021 · 41 citations
