VideoFusion: A Spatio-Temporal Collaborative Network for Multi-modal Video Fusion
Linfeng Tang, Yeda Wang, Meiqi Gong, Zizhuo Li, Yuxin Deng, Xunpeng Yi, Chunyu Li, Han Xu, Hao Zhang, Jiayi Ma
Abstract
Compared to images, videos better reflect real-world acquisition and possess valuable temporal cues. However, existing multi-sensor fusion research predominantly integrates complementary context from multiple images rather than videos due to the scarcity of large-scale multi-sensor video datasets, limiting research in video fusion and the inherent difficulty of jointly modeling spatial and temporal dependencies in a unified framework. To this end, we construct M3SVD, a benchmark dataset with 220 temporally synchronized and spatially registered infrared-visible videos comprising 153, 797 frames, bridging the data gap. Secondly, we propose VideoFusion, a multi-modal video fusion model that exploits cross-modal complementarity and temporal dynamics to generate spatio-temporally coherent videos from multi-modal inputs. Specifically, 1) a differential reinforcement module is developed for cross-modal information interaction and enhancement, 2) a complete modalityguided fusion strategy is employed to adaptively integrate multi-modal features, and 3) a bi-temporal co-attention mechanism is devised to dynamically aggregate forwardbackward temporal contexts to reinforce cross-frame feature representations. Experiments reveal that VideoFusion outperforms existing image-oriented fusion paradigms in sequences, effectively mitigating temporal inconsistency and interference. Project and M3SVD: https: //github.com/Linfeng-Tang/VideoFusion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2e1c845-dc86-42f1-87b7-d57e1e08b8f7Cited by top-tier papers4
- A Unified Solution to Video Fusion: From Multi-Frame Learning to BenchmarkingZixiang Zhao, Haowen Bai, Bingxin Ke, Yukun Cui et al.NeurIPS 2025 · 21 citations
- Multi-Modal Image Fusion via Intervention-Stable Feature LearningXue Wang, Zheng Guan, Wenhua Qian, Chengchao Wang et al.CVPR 2026 · 3 citations
- Streaming Diffusion Model for Fast Infrared and Visible Video FusionJinyuan Liu, Ludan Sun, Tengyu Ma, Chunyan Yang et al.CVPR 2026 · 2 citations
- AerialFusion: Co-Motion-Driven Unified Registration and Fusion on Multi-modal Data Streams from Aerial ViewJunhui Qiu, Xiang Xiang, Hongyun Wang, Jiaqi GuiAAAI 2026
Builds on20
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat et al.CVPR 2022 · 3,348 citations
- Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object DetectionJinyuan Liu, Xin Fan, Zhanbo Huang, Guanyao Wu et al.CVPR 2022 · 929 citations
- Rethinking the Image Fusion: A Fast Unified Image Fusion Network based on Proportional Maintenance of Gradient and IntensityHao Zhang, Han Xu, Yang Xiao, Xiaojie Guo et al.AAAI 2020 · 583 citations
- DDFM: Denoising Diffusion Model for Multi-Modality Image FusionZixiang Zhao, Haowen Bai, Yuanzhi Zhu, Jiangshe Zhang et al.ICCV 2023 · 350 citations
Related papers
- Multi-modal Gated Mixture of Local-to-Global Experts for Dynamic Image FusionBing Cao, Yiming Sun, Pengfei Zhu, Qinghua HuICCV 2023 · 110 citations
- TemCoCo: Temporally Consistent Multi-Modal Video Fusion with Visual-Semantic CollaborationMeiqi Gong, Hao Zhang, Xunpeng Yi, Linfeng Tang et al.ICCV 2025 · 4 citations
- OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video GenerationKepan Nan, Rui Xie, Penghao Zhou, Tiehan Fan et al.ICLR 2025
- ImViD: Immersive Volumetric Videos for Enhanced VR EngagementZhengxian Yang, Shi Pan, Shengqi Wang, Haoxiang Wang et al.CVPR 2025
- Bridging Human Evaluation to Infrared and Visible Image FusionJinyuan Liu, Xingyuan Li, Qingyun Mei, HaoYuan Xu et al.CVPR 2026 · 4 citations
