Streaming Diffusion Model for Fast Infrared and Visible Video Fusion
Jinyuan Liu, Ludan Sun, Tengyu Ma, Chunyan Yang, Zhiying Jiang, Long Ma, Risheng Liu, Xin Fan
摘要
Infrared and visible video fusion is pivotal for robust perceptual systems, aiming to synthesize a comprehensive video stream that leverages both thermal resilience and textured details. However, prevailing methods, by treating videos as sequences of independent frames, inherently introduce temporal incoherence, such as flickering and ghosting artifacts. While diffusion models possess strong generative priors to remedy this, their iterative nature is prohibitively slow for video. To resolve this fundamental dilemma, we propose a streaming diffusion model for efficient infrared and visible video fusion, termed SDMFusion. Our key insight is to exploit the generative prior of a pre-trained diffusion model into a one-step sampling framework, while explicitly modeling temporal dynamics. We design a memory-augmented latent pipeline where a temporal aggregation adapter aligns and propagates crossframe features to ensure coherence, supported by a dedicated temporal consistency loss. This approach effectively decouples the challenge of achieving high fidelity from maintaining temporal stability. Extensive experiments on four benchmarks demonstrate that our method establishes a new state-of-the-art, generating fused videos with exceptional spatio-temporal consistency at a speed suitable for real-time application. The code is available at https://github.com/DandanYoung/SDMFusion.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object DetectionJinyuan Liu, Xin Fan, Zhanbo Huang, Guanyao Wu 等CVPR 2022 · 被引用 929 次
- DDFM: Denoising Diffusion Model for Multi-Modality Image FusionZixiang Zhao, Haowen Bai, Yuanzhi Zhu, Jiangshe Zhang 等ICCV 2023 · 被引用 350 次
- Multi-interactive Feature Learning and a Full-time Multi-modality Benchmark for Image Fusion and SegmentationJinyuan Liu, Zhu Liu, Guanyao Wu, Long Ma 等ICCV 2023 · 被引用 287 次
- Equivariant Multi-Modality Image FusionZixiang Zhao, Haowen Bai, Jiangshe Zhang, Yulun Zhang 等CVPR 2024 · 被引用 155 次
- Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image FusionXunpeng Yi, Han Xu, Hao Zhang, Linfeng Tang 等CVPR 2024 · 被引用 121 次
相关 Paper
- DRFusion: Drift-Resilient Temporally Consistent Infrared–Visible Video FusionXingyuan Li, HaoYuan Xu, Shulin Li, Xiang Chen 等ICML 2026
- DRMF: Degradation-Robust Multi-Modal Image Fusion via Composable Diffusion PriorLinfeng Tang, Yuxin Deng, Xunpeng Yi, Qinglong Yan 等ACM MM 2024 · 被引用 53 次
- InstantViR: Real-Time Video Inverse Problem Solver with Distilled Diffusion PriorWeimin Bai, Suzhe Xu, Yiwei Ren, Jinhua Hao 等CVPR 2026 · 被引用 3 次
- Thermal Diffusion Matters: Infrared Spatial-Temporal Video Super-Resolution through Heat Conduction PriorsMingxuan Zhou, Shuang Li, Yutang Zhang, Jing Geng 等CVPR 2026
- STDD: Spatio-Temporal Dual Diffusion for Video GenerationShuaizhen Yao, Xiaoya Zhang, Xin Liu, Mengyi Liu 等CVPR 2025
