Every SAM Drop Counts: Embracing Semantic Priors for Multi-Modality Image Fusion and Beyond
Guanyao Wu, Haoyu Liu, Hongming Fu, Yichuan Peng, Jinyuan Liu, Xin Fan, Risheng Liu
Abstract
Multi-modality image fusion, particularly infrared and visible, plays a crucial role in integrating diverse modalities to enhance scene understanding. Although early research prioritized visual quality, preserving fine details and adapting to downstream tasks remains challenging. Recent approaches attempt task-specific design but rarely achieve "The Best of Both Worlds" due to inconsistent optimization goals. To address these issues, we propose a novel method that leverages the semantic knowledge from the Segment Anything Model (SAM) to Grow the quality of fusion results and Enable downstream task adaptability, namely SAGE. Specifically, we design a Semantic Persistent Attention (SPA) Module that efficiently maintains source information via the persistent repository while extracting high-level semantic priors from SAM. More importantly, to eliminate the impractical dependence on SAM during inference, we introduce a bi-level optimization-driven distillation mechanism with triplet losses, which allow the student network to effectively extract knowledge. Extensive experiments show that our method achieves a balance between high-quality visual results and downstream task adaptability while maintaining practical deployment efficiency. The code is available at https://github . com/RollingPlain/SAGE_IVIF.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0ab935a8-aadc-47d5-874f-59b6f9d850daCited by top-tier papers21
- Customized Fusion: A Closed-Loop Dynamic Network for Adaptive Multi-Task-Aware Infrared-Visible Image FusionZengyi Yang, Yu Liu, Juan Cheng, Zhiqin Zhu et al.CVPR 2026 · 7 citations
- Text-Guided Channel Perturbation and Pre-Trained Knowledge Integration for Unified Multi-Modality Image FusionXilai Li, Xiaosong Li, Weijun JiangAAAI 2026 · 4 citations
- Image Stitching in Adverse Condition: A Bidirectional-Consistency Learning Framework and BenchmarkZengxi Zhang, Junchen Ge, Zhiying Jiang, Miao Zhang et al.NeurIPS 2025 · 3 citations
- Missing No More: Dictionary-Guided Cross-Modal Image Fusion under Missing InfraredYafei Zhang, Meng Ma, Huafeng Li, Yu LiuCVPR 2026 · 3 citations
- Multi-Modal Image Fusion via Intervention-Stable Feature LearningXue Wang, Zheng Guan, Wenhua Qian, Chengchao Wang et al.CVPR 2026 · 3 citations
Builds on17
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object DetectionJinyuan Liu, Xin Fan, Zhanbo Huang, Guanyao Wu et al.CVPR 2022 · 929 citations
- Toward Fast, Flexible, and Robust Low-Light Image EnhancementLong Ma, Tengyu Ma, Risheng Liu, Xin Fan et al.CVPR 2022 · 928 citations
- Segment Everything Everywhere All at OnceXueyan Zou, Jianwei Yang, Hao Zhang, Feng Li et al.NeurIPS 2023 · 889 citations
- DDFM: Denoising Diffusion Model for Multi-Modality Image FusionZixiang Zhao, Haowen Bai, Yuanzhi Zhu, Jiangshe Zhang et al.ICCV 2023 · 350 citations
Related papers
- Distilling Semantic Priors from SAM to Efficient Image Restoration ModelsQuan Zhang, Xiaoyu Liu, Wei Li, Hanting Chen et al.CVPR 2024 · 20 citations
- Unleashing the Power of Generic Segmentation Model: A Simple Baseline for Infrared Small Target DetectionMingjin Zhang, Chi Zhang, Qiming Zhang, Yunsong Li et al.ACM MM 2024 · 33 citations
- SAM-Guided Semantic Knowledge Fusion for Visible-Infrared Object DetectionTing Li, Songtao Li, Shuaifeng Li, Xiaolin Qin et al.ACM MM 2025 · 2 citations
- Improving SAM for Camouflaged Object Detection via Dual Stream AdaptersJiaming Liu, Linghe Kong, Guihai ChenICCV 2025 · 5 citations
- Segment Anything Model Meets Semi-supervised Medical Image Segmentation: A Novel PerspectiveHaifeng Zhao, Haiyang Li, Lei-Lei Ma, Dengdi SunNeurIPS 2025 · 1 citation
