DRFusion: Drift-Resilient Temporally Consistent Infrared–Visible Video Fusion
Xingyuan Li, HaoYuan Xu, Shulin Li, Xiang Chen, Zhiying Jiang, Jinyuan Liu
Abstract
Infrared and visible video fusion is essential for achieving comprehensive perception in dynamic scenes. However, maintaining temporal consistency remains a formidable challenge. Conventional methods relying on optical flow often suffer from geometric rigidity and ghosting artifacts. Moreover, standard diffusion-based fusion models typically operate in a frame-by-frame manner; when extended to autoregressive settings, they lack intrinsic temporal constraints and are prone to severe error accumulation and drifting, where minor artifacts amplify over time. To address these limitations, we propose a drift-resilient video fusion method that reformulates the task as history-conditioned motion generation. We introduce Stabilized History Guidance and Soft Temporal Anchoring to reframe temporal consistency as spectral filtering, implicitly aggregating motion dynamics without rigid alignment. Furthermore, our Decoupled Structure-Motion Adaptation strategy bridges pre-trained priors and structural constraints via two-stage training and latent refinement. Extensive experiments demonstrate that our method achieves state-of-the-art performance in both fusion quality and temporal stability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f706bc29-7516-49de-adfa-b5c99a000804Builds on24
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
- Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video GenerationJay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei et al.ICCV 2023 · 1,113 citations
- Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object DetectionJinyuan Liu, Xin Fan, Zhanbo Huang, Guanyao Wu et al.CVPR 2022 · 929 citations
- DDFM: Denoising Diffusion Model for Multi-Modality Image FusionZixiang Zhao, Haowen Bai, Yuanzhi Zhu, Jiangshe Zhang et al.ICCV 2023 · 350 citations
Related papers
- Streaming Diffusion Model for Fast Infrared and Visible Video FusionJinyuan Liu, Ludan Sun, Tengyu Ma, Chunyan Yang et al.CVPR 2026 · 2 citations
- MV-Diffusion: Motion-aware Video Diffusion ModelZijun Deng, Xiangteng He, Yuxin Peng, Xiongwei Zhu et al.ACM MM 2023 · 20 citations
- WorldWeaver: Generating Long-Horizon Video Worlds via Rich PerceptionZhiheng Liu, Xueqing Deng, Shoufa Chen, Angtian Wang et al.NeurIPS 2025 · 15 citations
- STCDiT: Spatio-Temporally Consistent Diffusion Transformer for High-Quality Video Super-ResolutionJunyang Chen, Jiangxin Dong, Long Sun, Yixin Yang et al.CVPR 2026 · 1 citation
- Trajectory-Stabilized Inference for Diffusion-Based Video InpaintingZhanhe Zhang, Jiahua Li, Xu Yang, Kun Wei et al.ICML 2026
