Audio-Driven Identity Manipulation for Face Inpainting
Yuqi Sun, Qing Lin, Weimin Tan, Bo Yan
摘要
Recent advances in multimodal artificial intelligence have greatly improved the integration of vision-language-audio cues to enrich the content creation process. Inspired by these developments, in this paper, we first integrate audio into the face inpainting task to facilitate identity manipulation. Our main insight is that a person's voice carries distinct identity markers, such as age and gender, which provide an essential supplement for identity-aware face inpainting. By extracting identity information from audio as guidance, our method can naturally support tasks of identity preservation and identity swapping in face inpainting. Specifically, we introduce a dual-stream network architecture comprising a face branch and an audio branch. The face branch is tasked with extracting deterministic information from the visible parts of the input masked face, while the audio branch is designed to capture heuristic identity priors from the speaker's voice. The identity codes from two streams are integrated using a multi-layer perceptron (MLP) to create a virtual unified identity embedding that represennts comprehensive identity features. In addition, to explicitly exploit the information from audio, we introduce an audio-face generator to generate an 'fake' audio face directly from audio and fuse the multi-scale intermediate features from the audio-face generator into face inpainting network through an audio-visual feature fusion (AVFF) module. Extensive experiments demonstrate the positive impact of extracting identity information from audio on face inpainting task, especially in identity preservation.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- One-stage Context and Identity Hallucination NetworkYinglu Liu, Mingcan Xiang, Hailin Shi, Tao MeiACM MM 2021 · 被引用 2 次
- Learning to Have an Ear for Face Super-ResolutionGivi Meishvili, Simon Jenni, Paolo FavaroCVPR 2020
- From Inference to Generation: End-to-end Fully Self-supervised Generation of Human Face from SpeechHyeong-Seok Choi, Changdae Park, Kyogu LeeICLR 2020 · 被引用 33 次
- FlowFace: Semantic Flow-Guided Shape-Aware Face SwappingHao Zeng, Wei Zhang, Changjie Fan, Tangjie Lv 等AAAI 2023 · 被引用 11 次
- Identity-Preserving Talking Face Generation with Landmark and Appearance PriorsWeizhi Zhong, Chaowei Fang, Yinqi Cai, Pengxu Wei 等CVPR 2023
