VideoFusion: A Spatio-Temporal Collaborative Network for Multi-modal Video Fusion
Linfeng Tang, Yeda Wang, Meiqi Gong, Zizhuo Li, Yuxin Deng, Xunpeng Yi, Chunyu Li, Han Xu, Hao Zhang, Jiayi Ma
摘要
Compared to images, videos better reflect real-world acquisition and possess valuable temporal cues. However, existing multi-sensor fusion research predominantly integrates complementary context from multiple images rather than videos due to the scarcity of large-scale multi-sensor video datasets, limiting research in video fusion and the inherent difficulty of jointly modeling spatial and temporal dependencies in a unified framework. To this end, we construct M3SVD, a benchmark dataset with 220 temporally synchronized and spatially registered infrared-visible videos comprising 153, 797 frames, bridging the data gap. Secondly, we propose VideoFusion, a multi-modal video fusion model that exploits cross-modal complementarity and temporal dynamics to generate spatio-temporally coherent videos from multi-modal inputs. Specifically, 1) a differential reinforcement module is developed for cross-modal information interaction and enhancement, 2) a complete modalityguided fusion strategy is employed to adaptively integrate multi-modal features, and 3) a bi-temporal co-attention mechanism is devised to dynamically aggregate forwardbackward temporal contexts to reinforce cross-frame feature representations. Experiments reveal that VideoFusion outperforms existing image-oriented fusion paradigms in sequences, effectively mitigating temporal inconsistency and interference. Project and M3SVD: https: //github.com/Linfeng-Tang/VideoFusion.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- A Unified Solution to Video Fusion: From Multi-Frame Learning to BenchmarkingZixiang Zhao, Haowen Bai, Bingxin Ke, Yukun Cui 等NeurIPS 2025 · 被引用 21 次
- Multi-Modal Image Fusion via Intervention-Stable Feature LearningXue Wang, Zheng Guan, Wenhua Qian, Chengchao Wang 等CVPR 2026 · 被引用 3 次
- Streaming Diffusion Model for Fast Infrared and Visible Video FusionJinyuan Liu, Ludan Sun, Tengyu Ma, Chunyan Yang 等CVPR 2026 · 被引用 2 次
- AerialFusion: Co-Motion-Driven Unified Registration and Fusion on Multi-modal Data Streams from Aerial ViewJunhui Qiu, Xiang Xiang, Hongyun Wang, Jiaqi GuiAAAI 2026
它引用的顶会 Paper20
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat 等CVPR 2022 · 被引用 3,348 次
- Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object DetectionJinyuan Liu, Xin Fan, Zhanbo Huang, Guanyao Wu 等CVPR 2022 · 被引用 929 次
- Rethinking the Image Fusion: A Fast Unified Image Fusion Network based on Proportional Maintenance of Gradient and IntensityHao Zhang, Han Xu, Yang Xiao, Xiaojie Guo 等AAAI 2020 · 被引用 583 次
- DDFM: Denoising Diffusion Model for Multi-Modality Image FusionZixiang Zhao, Haowen Bai, Yuanzhi Zhu, Jiangshe Zhang 等ICCV 2023 · 被引用 350 次
相关 Paper
- Multi-modal Gated Mixture of Local-to-Global Experts for Dynamic Image FusionBing Cao, Yiming Sun, Pengfei Zhu, Qinghua HuICCV 2023 · 被引用 110 次
- TemCoCo: Temporally Consistent Multi-Modal Video Fusion with Visual-Semantic CollaborationMeiqi Gong, Hao Zhang, Xunpeng Yi, Linfeng Tang 等ICCV 2025 · 被引用 4 次
- OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video GenerationKepan Nan, Rui Xie, Penghao Zhou, Tiehan Fan 等ICLR 2025
- ImViD: Immersive Volumetric Videos for Enhanced VR EngagementZhengxian Yang, Shi Pan, Shengqi Wang, Haoxiang Wang 等CVPR 2025
- Bridging Human Evaluation to Infrared and Visible Image FusionJinyuan Liu, Xingyuan Li, Qingyun Mei, HaoYuan Xu 等CVPR 2026 · 被引用 4 次
