Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality
Zekai Luo, Zongze Du, Zhouhang Zhu, Hao Zhong, Muzhi Zhu, Wen Wang, Yuling Xi, Chenchen Jing, Hao Chen, Chunhua Shen
Abstract
Video face swapping is crucial in film and entertainment production, where achieving high fidelity and temporal consistency over long and complex video sequences remains a significant challenge. Inspired by recent advances in reference-guided image editing, we explore whether rich visual attributes from source videos can be similarly leveraged to enhance both fidelity and temporal coherence in video face swapping. Building on this insight, this work presents LivingSwap, the first video reference guided face swapping model. Our approach employs keyframes as conditioning signals to inject the target identity, enabling flexible and controllable editing. By combining keyframe conditioning with video reference guidance, the model performs temporal stitching to ensure stable identity preservation and high-fidelity reconstruction across long video sequences. To address the scarcity of data for reference-guided training, we construct a paired face-swapping dataset, Face2Face, and further reverse the data pairs to ensure reliable ground-truth supervision. Extensive experiments demonstrate that our method achieves state-of-the-art results, seamlessly integrating the target identity with the source video's expressions, lighting, and motion, while significantly reducing manual effort in production workflows. Project webpage: https://aim-uofa.github.io/LivingSwap
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 45c506e8-1a12-48ae-ae41-933972ecd00aBuilds on28
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- FaceForensics++: Learning to Detect Manipulated Facial ImagesAndreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess et al.ICCV 2019 · 2,966 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and EditingMingdeng Cao, Xintao Wang, Zhongang Qi, Ying Shan et al.ICCV 2023 · 770 citations
Related papers
- VividFace: A Robost and High-Fidelity Video Face Swapping FrameworkHao Shao, Shulun Wang, Yang Zhou, Guanglu Song et al.NeurIPS 2025 · 4 citations
- Canonswap: High-Fidelity and Consistent Video Face Swapping Via Canonical Space ModulationXiangyang Luo, Ye Zhu, Yunfei Liu, Lijian Lin et al.ICCV 2025 · 4 citations
- DynamicFace: High-Quality and Consistent Face Swapping for Image and Video Using Composable 3D Facial PriorsRunqi Wang, Yang Chen, Sijie Xu, Tianyao He et al.ICCV 2025 · 7 citations
- DreamSwapV: Mask-guided Subject Swapping for Any Customized Video EditingWeitao Wang, Zichen Wang, Hongdeng Shen, Yulei Lu et al.ICLR 2026 · 1 citation
- High-Fidelity Diffusion Face Swapping with ID-Constrained Facial ConditioningDailan He, Xiahong Wang, Shulun Wang, Hao Shao et al.CVPR 2026 · 5 citations
