Learning Adaptive Spatial Coherent Correlations for Speech-Preserving Facial Expression Manipulation
Tianshui Chen, Jianman Lin, Zhijing Yang, Chunmei Qing, Liang Lin
摘要
Speech-preserving facial expression manipulation (SPFEM) aims to modify facial emotions while meticulously maintaining the mouth animation associated with spoken content. Current works depend on inaccessible paired training samples for the person, where two aligned frames exhibit the same speech content yet differ in emotional expression, limiting the SPFEM applications in real-world scenarios. In this work, we discover that speakers who convey the same content with different emotions exhibit highly correlated local facial animations, providing valuable supervision for SPFEM. To capitalize on this insight, we propose a novel adaptive spatial coherent correlation learning (ASCCL) algorithm, which models the aforementioned correlation as an explicit metric and integrates the metric to supervise manipulating facial expression and meanwhile better preserving the facial animation of spoken contents. To this end, it first learns a spatial coherent correlation metric, ensuring the visual disparities of adjacent local regions of the image belonging to one emotion are similar to those of the corresponding counterpart of the image belonging to another emotion. Recognizing that visual disparities are not uniform across all regions, we have also crafted a disparity-aware adaptive strategy that prioritizes regions that present greater challenges. During SPFEM model training, we construct the adaptive spatial coherent correlation metric between corresponding local regions of the input and output images as addition loss to supervise the generation * Zhijing Yang is the corresponding author. Tianshui Chen, Jianman Lin, and Zhijing Yang are with
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper10
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Alias-Free Generative Adversarial NetworksTero Karras, Miika Aittala, Samuli Laine, Erik Härkönen 等NeurIPS 2021 · 被引用 2,126 次
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 被引用 869 次
- Learning an animatable detailed 3D face model from in-the-wild imagesYao Feng, Haiwen Feng, Michael J. Black, Timo BolkartSIGGRAPH 2021 · 被引用 662 次
- Neural Emotion Director: Speech-preserving semantic control of facial expressions in "in-the-wild" videosFoivos Paraperas Papantoniou, Panagiotis Paraskevas Filntisis, Petros Maragos, Anastasios RoussosCVPR 2022 · 被引用 32 次
相关 Paper
- Self-Supervised Emotion Representation Disentanglement for Speech-Preserving Facial Expression ManipulationZhihua Xu, Tianshui Chen, Zhijing Yang, Chunmei Qing 等ACM MM 2024 · 被引用 5 次
- JDMAN: Joint Discriminative and Mutual Adaptation Networks for Cross-Domain Facial Expression RecognitionYingjian Li, Yingnan Gao, Bingzhi Chen, Zheng Zhang 等ACM MM 2021 · 被引用 21 次
- Learning Motion Refinement for Unsupervised Face AnimationJiale Tao, Shuhang Gu, Wen Li, Lixin DuanNeurIPS 2023 · 被引用 10 次
- BHGap: A Deep Iterative Prompting and Multi-stage Alignment Framework for Dynamic Facial Expression RecognitionYichi Zhang, Yunqi Han, Jiayue Ding, Liangyu ChenWWW 2026
- Contrastive Adversarial Learning for Person Independent Facial Emotion RecognitionDae Ha Kim, Byung Cheol SongAAAI 2021 · 被引用 41 次
