Video Harmonization with Triplet Spatio-Temporal Variation Patterns
Zonghui Guo, Xinyu Han, Jie Zhang, Shiguang Shan, Haiyong Zheng
摘要
Video harmonization is an important and challenging task that aims to obtain visually realistic composite videos by automatically adjusting the foreground's appearance to harmonize with the background. Inspired by the short-term and long-term gradual adjustment process of manual harmonization, we present a Video Triplet Transformer framework to model three spatio-temporal variation patterns within videos, i.e., short-term spatial as well as long-term global and dynamic, for video-to-video tasks like video harmonization. Specifically, for short-term harmonization, we adjust foreground appearance to consist with background in spatial dimension based on the neighbor frames; for long-term harmonization, we not only explore global appearance variations to enhance temporal consistency but also alleviate motion offset constraints to align similar contextual appearances dynamically. Extensive experiments and ablation studies demonstrate the effectiveness of our method, achieving state-of-the-art performance in video harmonization, video enhancement, and video demoiréing tasks. We also propose a temporal consistency metric to better evaluate the harmonized videos. Code is available at https://github.com/zhenglab/VideoTripletTransformer .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- DiffusionHarmonizer: Bridging Neural Reconstruction and Photorealistic Simulation with Online Diffusion EnhancerYuxuan Zhang, Katarína Tóthová, Zian Wang, Kangxue Yin 等CVPR 2026 · 被引用 10 次
- GenCompositor: Generative Video Compositing with Diffusion TransformerShuzhou Yang, Xiaoyu Li, Xiaodong Cun, Guangzhi Wang 等ICLR 2026 · 被引用 10 次
- Face Forgery Video Detection via Temporal Forgery Cue UnravelingZonghui Guo, Yingjie Liu, Jie Zhang, Haiyong Zheng 等CVPR 2025
它引用的顶会 Paper20
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu 等ICML 2020 · 被引用 1,773 次
相关 Paper
- Image Harmonization with TransformerZonghui Guo, Dongsheng Guo, Haiyong Zheng, Zhaorui Gu 等ICCV 2021 · 被引用 95 次
- Relightful Video Portrait HarmonizationJun Myeong Choi, Jae Shin Yoon, Luchao Qi, Roni Sengupta 等CVPR 2026
- Video Frame Interpolation TransformerZhihao Shi, Xiangyu Xu, Xiaohong Liu, Jun Chen 等CVPR 2022 · 被引用 117 次
- Decouple and Track: Benchmarking and Improving Video Diffusion Transformers for Motion TransferQingyu Shi, Jianzong Wu, Jinbin Bai, Jiangning Zhang 等ICCV 2025 · 被引用 1 次
- Learning Global-aware Kernel for Image HarmonizationXintian Shen, Jiangning Zhang, Jun Chen, Shipeng Bai 等ICCV 2023 · 被引用 14 次
