Probing Synergistic High-Order Interaction in Infrared and Visible Image Fusion
Naishan Zheng, Man Zhou, Jie Huang, Junming Hou, Haoying Li, Yuan Xu, Feng Zhao
Abstract
Infrared and visible image fusion aims to generate a fused image by integrating and distinguishing complementary information from multiple sources. While the cross-attention mechanism with global spatial interactions appears promising, it only capture second-order spatial inter-actions, neglecting higher-order interactions in both spatial and channel dimensions. This limitation hampers the ex-ploitation of synergies between multi-modalities. To bridge this gap, we introduce a Synergistic High-order Interaction Paradigm (SHIP), designed to systematically investigate the spatial fine-grained and global statistics collaborations between infrared and visible images across two fundamental dimensions: 1) Spatial dimension: we construct spatial fine-grained interactions through element-wise multiplication, mathematically equivalent to global interactions, and then foster high-order formats by iteratively aggregating and evolving complementary information, enhancing both efficiency andflexibility; 2) Channel dimension: expanding on channel interactions with first-order statistics (mean), we devise high-order channel interactions to facilitate the discernment of inter-dependencies between source images based on global statistics. Harnessing high-order interactions significantly enhances our model's ability to exploit multi-modal synergies, leading to superior performance over state-of-the-art alternatives, as shown through comprehensive experiments across various benchmarks. Code is available at https://github.com/zheng980629/SHIP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ac0054d6-551f-4a89-b0fa-76983d0e23d5Cited by top-tier papers9
- Residual Prior-driven Frequency-aware Network for Image FusionZheng Guan, Xue Wang, Wenhua Qian, Peng Liu et al.ACM MM 2025 · 10 citations
- Bridging Human Evaluation to Infrared and Visible Image FusionJinyuan Liu, Xingyuan Li, Qingyun Mei, HaoYuan Xu et al.CVPR 2026 · 4 citations
- 3M-TI: High-Quality Mobile Thermal Imaging via Calibration-free Multi-Camera Cross-Modal DiffusionMinchong Chen, Xiaoyun Yuan, Junzhe Wan, Jianing Zhang et al.CVPR 2026 · 2 citations
- RegionFuse: Region-Adaptive Pixel Distribution Learning for Infrared and Visible Image FusionJianghan Xia, Hong Song, Jinfu Li, Yucong Lin et al.CVPR 2026 · 1 citation
- VMDiff: Visual Mixing Diffusion for Limitless Cross-Object SynthesisZeren Xiong, Yue Yu, Ze-dong Zhang, Shuo Chen et al.ICLR 2026 · 1 citation
Builds on27
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object DetectionJinyuan Liu, Xin Fan, Zhanbo Huang, Guanyao Wu et al.CVPR 2022 · 929 citations
- Toward Fast, Flexible, and Robust Low-Light Image EnhancementLong Ma, Tengyu Ma, Risheng Liu, Xin Fan et al.CVPR 2022 · 928 citations
- Rethinking the Image Fusion: A Fast Unified Image Fusion Network based on Proportional Maintenance of Gradient and IntensityHao Zhang, Han Xu, Yang Xiao, Xiaojie Guo et al.AAAI 2020 · 583 citations
Related papers
- The Source Image Is the Best Attention for Infrared and Visible Image FusionSong Wang, Xie Han, Liqun Kuang, Boying Wang et al.ICCV 2025 · 6 citations
- Learning a Graph Neural Network with Cross Modality Interaction for Image FusionJiawei Li, Jiansheng Chen, Jinyuan Liu, Huimin MaACM MM 2023 · 85 citations
- Multi-modal Gated Mixture of Local-to-Global Experts for Dynamic Image FusionBing Cao, Yiming Sun, Pengfei Zhu, Qinghua HuICCV 2023 · 110 citations
- TeRF: Text-driven and Region-aware Flexible Visible and Infrared Image FusionHebaixu Wang, Hao Zhang, Xunpeng Yi, Xinyu Xiang et al.ACM MM 2024 · 11 citations
- DetFusion: A Detection-driven Infrared and Visible Image Fusion NetworkYiming Sun, Bing Cao, Pengfei Zhu, Qinghua HuACM MM 2022 · 165 citations
