Zero-Shot Head Swapping in Real-World Scenarios
Taewoong Kang, Sohyun Jeong, Hyojin Jang, Jaegul Choo
Abstract
With growing demand in media and social networks for personalized images, the need for advanced head-swapping techniques—integrating an entire head from the head image with the body from the body image—has increased. However, traditional head-swapping methods heavily rely on face-centered cropped data with primarily frontal-facing views, which limits their effectiveness in real-world applications. Additionally, their masking methods, designed to indicate regions requiring editing, are optimized for these types of dataset but struggle to achieve seamless blending in complex situations, such as when the original data includes features like long hair extending beyond the masked area. To overcome these limitations and enhance adaptability in diverse and complex scenarios, we propose a novel head swapping method, HID, that is robust to images including the full head and the upper body, and handles from frontal to side views, while automatically generating context-aware masks. For automatic mask generation, we introduce the IOMask, which enables seamless blending of the head and body, effectively addressing integration challenges. We further introduce the hair injection module to capture hair details with greater precision. Our experiments demonstrate that the proposed approach achieves state-of-the-art performance in head swapping, providing visually consistent and realistic results across a wide range of challenging conditions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Few-Shot Head Swapping in the WildChangyong Shu, Hemao Wu, Hang Zhou, Jiaming Liu et al.CVPR 2022 · 18 citations
- Controllable and Expressive One-Shot Video Head SwappingChaonan Ji, Jinwei Qi, Peng Zhang, Bang Zhang et al.ICCV 2025
- One-stage Context and Identity Hallucination NetworkYinglu Liu, Mingcan Xiang, Hailin Shi, Tao MeiACM MM 2021 · 2 citations
- DreamSwapV: Mask-guided Subject Swapping for Any Customized Video EditingWeitao Wang, Zichen Wang, Hongdeng Shen, Yulei Lu et al.ICLR 2026 · 1 citation
- HairCUP: Hair Compositional Universal Prior for 3D Gaussian AvatarsByungjun Kim, Shunsuke Saito, Giljoo Nam, Tomas Simon et al.ICCV 2025 · 2 citations
