Video Harmonization with Triplet Spatio-Temporal Variation Patterns
Zonghui Guo, Xinyu Han, Jie Zhang, Shiguang Shan, Haiyong Zheng
Abstract
Video harmonization is an important and challenging task that aims to obtain visually realistic composite videos by automatically adjusting the foreground's appearance to harmonize with the background. Inspired by the short-term and long-term gradual adjustment process of manual harmonization, we present a Video Triplet Transformer framework to model three spatio-temporal variation patterns within videos, i.e., short-term spatial as well as long-term global and dynamic, for video-to-video tasks like video harmonization. Specifically, for short-term harmonization, we adjust foreground appearance to consist with background in spatial dimension based on the neighbor frames; for long-term harmonization, we not only explore global appearance variations to enhance temporal consistency but also alleviate motion offset constraints to align similar contextual appearances dynamically. Extensive experiments and ablation studies demonstrate the effectiveness of our method, achieving state-of-the-art performance in video harmonization, video enhancement, and video demoiréing tasks. We also propose a temporal consistency metric to better evaluate the harmonized videos. Code is available at https://github.com/zhenglab/VideoTripletTransformer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bf71b5dc-fe46-4b2a-a9c1-a6a8227cae22Cited by top-tier papers3
- DiffusionHarmonizer: Bridging Neural Reconstruction and Photorealistic Simulation with Online Diffusion EnhancerYuxuan Zhang, Katarína Tóthová, Zian Wang, Kangxue Yin et al.CVPR 2026 · 10 citations
- GenCompositor: Generative Video Compositing with Diffusion TransformerShuzhou Yang, Xiaoyu Li, Xiaodong Cun, Guangzhi Wang et al.ICLR 2026 · 10 citations
- Face Forgery Video Detection via Temporal Forgery Cue UnravelingZonghui Guo, Yingjie Liu, Jie Zhang, Haiyong Zheng et al.CVPR 2025
Builds on20
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
Related papers
- Image Harmonization with TransformerZonghui Guo, Dongsheng Guo, Haiyong Zheng, Zhaorui Gu et al.ICCV 2021 · 95 citations
- Relightful Video Portrait HarmonizationJun Myeong Choi, Jae Shin Yoon, Luchao Qi, Roni Sengupta et al.CVPR 2026
- Video Frame Interpolation TransformerZhihao Shi, Xiangyu Xu, Xiaohong Liu, Jun Chen et al.CVPR 2022 · 117 citations
- Decouple and Track: Benchmarking and Improving Video Diffusion Transformers for Motion TransferQingyu Shi, Jianzong Wu, Jinbin Bai, Jiangning Zhang et al.ICCV 2025 · 1 citation
- Learning Global-aware Kernel for Image HarmonizationXintian Shen, Jiangning Zhang, Jun Chen, Shipeng Bai et al.ICCV 2023 · 14 citations
