CoDeF: Content Deformation Fields for Temporally Consistent Video Processing
Hao Ouyang, Qiuyu Wang, Yuxi Xiao, Qingyan Bai, Juntao Zhang, Kecheng Zheng, Xiaowei Zhou, Qifeng Chen, Yujun Shen
Abstract
We present the content deformation field (CoDeF) as a new type of video representation, which consists of a canonical content field aggregating the static contents in the entire video and a temporal deformation field recording the transformations from the canonical image (i.e., rendered from the canonical content field) to each individual frame along the time axis. Given a target video, these two fields are jointly optimized to reconstruct it through a carefully tailored rendering pipeline. We advisedly introduce some regularizations into the optimization process, urging the canonical content field to inherit semantics (e.g., the object shape) from the video. With such a design, CoDeF naturally supports lifting image algorithms for video processing, in the sense that one can apply an image algorithm to the canonical image and effortlessly propagate the outcomes to the entire video with the aid of the temporal deformation field. We experimentally show that CoDeF is able to lift image-to-image translation to video-to-video translation and lift keypoint detection to keypoint tracking without any training. More importantly, thanks to our lifting strategy that deploys the algorithms on only one image, we achieve superior cross-frame consistency in processed videos compared to existing video-to-video translation approaches, and even manage to track non-rigid objects like water and smog. Code is made available at https: / /qiuyu96. github.io/CoDeF/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers47
- Scaling Instruction-Based Video Editing with a High-Quality Synthetic DatasetQingyan Bai, Qiuyu Wang, Hao Ouyang, Yue Yu et al.CVPR 2026 · 79 citations
- GFlow: Recovering 4D World from Monocular VideoShizun Wang, Xingyi Yang, Qiuhong Shen, Zhenxiang Jiang et al.AAAI 2025 · 47 citations
- EasyCreator: Empowering 4D Creation through Video InpaintingYue Ma, Kunyu Feng, Xinhua Zhang, Hongyu Liu et al.ICLR 2026 · 47 citations
- Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object MotionShiyuan Yang, Liang Hou, Haibin Huang, Chongyang Ma et al.SIGGRAPH 2024 · 46 citations
- Splatter a Video: Video Gaussian Representation for Versatile ProcessingYang-Tian Sun, Yihua Huang, Lin Ma, Xiaoyang Lyu et al.NeurIPS 2024 · 41 citations
Builds on37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- KeyTr: Keypoint Transporter for 3D Reconstruction of Deformable Objects in VideosDavid Novotný, Ignacio Rocco, Samarth Sinha, Alexandre Carlier et al.CVPR 2022 · 11 citations
- Generative Video Motion Editing with 3D Point TracksYao-Chih Lee, Zhoutong Zhang, Jiahui Huang, Jui-Hsien Wang et al.CVPR 2026 · 23 citations
- Neural Inter-Frame Compression for Video CodingAbdelaziz Djelouah, Joaquim Campos, Simone Schaub-Meyer, Christopher SchroersICCV 2019 · 207 citations
- STRIVE: Scene Text Replacement In VideosVijay Kumar B. G, Jeyasri Subramanian, Varnith Chordia, Eugene Bart et al.ICCV 2021 · 14 citations
- Deformable Sprites for Unsupervised Video DecompositionVickie Ye, Zhengqi Li, Richard Tucker, Angjoo Kanazawa et al.CVPR 2022 · 45 citations
