Delving into the Frequency: Temporally Consistent Human Motion Transfer in the Fourier Space
Guang Yang, Wu Liu, Xinchen Liu, Xiaoyan Gu, Juan Cao, Jintao Li
摘要
Human motion transfer refers to synthesizing photo-realistic and temporally coherent videos that enable one person to imitate the motion of others. However, current synthetic videos suffer from the temporal inconsistency in sequential frames that significantly degrades the video quality, yet is far from solved by existing methods in the pixel domain. Recently, some works on DeepFake detection try to distinguish the natural and synthetic images in the frequency domain because of the frequency insufficiency of image synthesizing methods. Nonetheless, there is no work to study the temporal inconsistency of synthetic videos from the aspects of the frequency-domain gap between natural and synthetic videos. Therefore, in this paper, we propose to delve into the frequency space for temporally consistent human motion transfer. First of all, we make the first comprehensive analysis of natural and synthetic videos in the frequency domain to reveal the frequency gap in both the spatial dimension of individual frames and the temporal dimension of the video. To close the frequency gap between the natural and synthetic videos, we propose a novel Frequency-based human MOtion TRansfer framework, named FreMOTR, which can effectively mitigate the spatial artifacts and the temporal inconsistency of the synthesized videos. FreMOTR explores two novel frequency-based regularization modules: 1) the Frequency-domain Appearance Regularization (FAR) to improve the appearance of the person in individual frames and 2) Temporal Frequency Regularization (TFR) to guarantee the temporal consistency between adjacent frames. Finally, comprehensive experiments demonstrate that the FreMOTR not only yields superior performance in temporal consistency metrics but also improves the frame-level visual quality of synthetic videos. In particular, the temporal consistency metrics are improved by nearly 30% than the state-of-the-art model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Spectral Motion Alignment for Video Motion Transfer Using Diffusion ModelsGeon Yeong Park, Hyeonho Jeong, Sang Wan Lee, Jong Chul YeAAAI 2025 · 被引用 21 次
- Decompose More and Aggregate Better: Two Closer Looks at Frequency Representation Learning for Human Motion PredictionXuehao Gao, Shaoyi Du, Yang Wu, Yang YangCVPR 2023
- Make-Your-Anchor: A Diffusion-based 2D Avatar Generation FrameworkZiyao Huang, Fan Tang, Yong Zhang, Xiaodong Cun 等CVPR 2024
它引用的顶会 Paper12
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil 等NeurIPS 2020 · 被引用 4,036 次
- Leveraging Frequency Analysis for Deep Fake Image RecognitionJoel Frank, Thorsten Eisenhofer, Lea Schönherr, Asja Fischer 等ICML 2020 · 被引用 848 次
- Fast Fourier ConvolutionLu Chi, Borui Jiang, Yadong MuNeurIPS 2020 · 被引用 842 次
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 被引用 840 次
- ClothFlow: A Flow-Based Model for Clothed Person GenerationXintong Han, Weilin Huang, Xiaojun Hu, Matthew R. ScottICCV 2019 · 被引用 297 次
相关 Paper
- Beyond Spatial Frequency: Pixel-Wise Temporal Frequency-Based Deepfake Video DetectionTaehoon Kim, Jongwook Choi, Yonghyun Jeong, Haeun Noh 等ICCV 2025 · 被引用 7 次
- C2F-FWN: Coarse-to-Fine Flow Warping Network for Spatial-Temporal Consistent Motion TransferDongxu Wei, Xiaowei Xu, Haibin Shen, Kejie HuangAAAI 2021 · 被引用 23 次
- REMOT: A Region-to-Whole Framework for Realistic Human Motion TransferQuanwei Yang, Xinchen Liu, Wu Liu, Hongtao Xie 等ACM MM 2022 · 被引用 5 次
- Frequency-Aware Spatiotemporal Transformers for Video Inpainting DetectionBingyao Yu, Wanhua Li, Xiu Li, Jiwen Lu 等ICCV 2021 · 被引用 38 次
- Spatio-Temporal Catcher: A Self-Supervised Transformer for Deepfake Video DetectionMaosen Li, Xurong Li, Kun Yu, Cheng Deng 等ACM MM 2023 · 被引用 9 次
