Diffusion Video Autoencoders: Toward Temporally Consistent Face Video Editing via Disentangled Video Encoding
Gyeongman Kim, Hajin Shim, Hyunsu Kim, Yunjey Choi, Junho Kim, Eunho Yang
摘要
Inspired by the impressive performance of recent face image editing methods, several studies have been naturally proposed to extend these methods to the face video editing task. One of the main challenges here is temporal consistency among edited frames, which is still unresolved. To this end, we propose a novel face video editing framework based on diffusion autoencoders that can successfully extract the decomposed features - for the first time as a face video editing model - of identity and motion from a given video. This modeling allows us to edit the video by simply manipulating the temporally invariant feature to the desired direction for the consistency. Another unique strength of our model is that, since our model is based on diffusion models, it can satisfy both reconstruction and edit capabilities at the same time, and is robust to corner cases in wild face videos (e.g. occluded faces) unlike the existing GAN-based methods.11Project page: https://diff-video-ae.github.io
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- FinePOSE: Fine-Grained Prompt-Driven 3D Human Pose Estimation via Diffusion ModelsJinglin Xu, Yijie Guo, Yuxin PengCVPR 2024 · 被引用 39 次
- DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion AutoencoderChenpeng Du, Qi Chen, Tianyu He, Xu Tan 等ACM MM 2023 · 被引用 36 次
- Language-driven Object Fusion into Neural Radiance Fields with Pose-Conditioned Dataset UpdatesKa-Chun Shum, Jaeyeon Kim, Binh-Son Hua, Duc Thanh Nguyen 等CVPR 2024 · 被引用 7 次
- DevFD : Developmental Face Forgery Detection by Learning Shared and Orthogonal LoRA SubspacesTianshuo Zhang, Li Gao, Siran Peng, Xiangyu Zhu 等NeurIPS 2025 · 被引用 4 次
- ConsistentAvatar: Learning to Diffuse Fully Consistent Talking Head Avatar with Temporal GuidanceHaijie Yang, Zhenyu Zhang, Hao Tang, Jianjun Qian 等ACM MM 2024 · 被引用 3 次
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 被引用 5,234 次
相关 Paper
- StableVideo: Text-driven Consistency-aware Diffusion Video EditingWenhao Chai, Xun Guo, Gaoang Wang, Yan LuICCV 2023 · 被引用 219 次
- TokenFlow: Consistent Diffusion Features for Consistent Video EditingMichal Geyer, Omer Bar-Tal, Shai Bagon, Tali DekelICLR 2024 · 被引用 439 次
- DeepFaceVideoEditing: sketch-based deep editing of face videosFeng-Lin Liu, Shu-Yu Chen, Yu-Kun Lai, Chunpeng Li 等SIGGRAPH 2022 · 被引用 26 次
- RIGID: Recurrent GAN Inversion and Editing of Real Face VideosYangyang Xu, Shengfeng He, Kwan-Yee K. Wong, Ping LuoICCV 2023 · 被引用 14 次
- DIVE: Taming DINO for Subject-Driven Video EditingYi Huang, Wei Xiong, He Zhang, Chaoqi Chen 等ICCV 2025 · 被引用 1 次
