Pixel-Perfect Puppetry: Precision-Guided Enhancement for Face Image and Video Editing
Yan Li, Zhenyi Wang, Guanghao Li, Wei Xue, Wenhan Luo, Yike Guo
Abstract
Preserving identity while precisely manipulating attributes is a central challenge in face editing for both images and videos. Existing methods often introduce visual artifacts or fail to maintain temporal consistency. We present FlowGuide, a unified framework that achieves fine-grained control over face editing in diffusion models. Our approach is founded on the local linearity of the UNet bottleneck’s latent space, which allows us to treat semantic attributes as corresponding to specific linear subspaces, providing a mathematically sound basis for disentanglement. FlowGuide first identifies a set of orthogonal basis vectors that span these semantic subspaces for both the original content and the target edit, a representation that efficiently captures the most salient features of each. We then introduce a novel guidance mechanism that quantifies the geometric alignment between these bases to dynamically steer the denoising trajectory at each step. This approach offers superior control by ensuring edits are confined to the desired attribute’s semantic axis while preserving orthogonal components related to identity. Extensive experiments demonstrate that FlowGuide achieves state-of-the-art performance, producing high-quality edits with superior identity preservation and temporal coherence. Our code is available at: https://github.com/yl4467/flow_edit.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
- GANSpace: Discovering Interpretable GAN ControlsErik Härkönen, Aaron Hertzmann, Jaakko Lehtinen, Sylvain ParisNeurIPS 2020 · 1,049 citations
Related papers
- Unsupervised Region-Based Image Editing of Denoising Diffusion ModelsZixiang Li, Yue Song, Renshuai Tao, Xiaohong Jia et al.AAAI 2025 · 1 citation
- DisControlFace: Adding Disentangled Control to Diffusion Autoencoder for One-shot Explicit Facial Image EditingHaozhe Jia, Yan Li, Hengfei Cui, Di Xu et al.ACM MM 2024 · 2 citations
- DeX-Portrait: Disentangled and Expressive Portrait Animation via Explicit and Latent Motion RepresentationsYuxiang Shi, Zhe Li, Yanwen Wang, Hao Zhu et al.CVPR 2026 · 3 citations
- High-Fidelity Diffusion Face Swapping with ID-Constrained Facial ConditioningDailan He, Xiahong Wang, Shulun Wang, Hao Shao et al.CVPR 2026 · 5 citations
- Continuous Control of Editing Models via Adaptive-Origin GuidanceAlon Wolf, Chen Katzir, Kfir Aberman, Or PatashnikSIGGRAPH 2026
