ControlFusion: A Controllable Image Fusion Network with Language-Vision Degradation Prompts
Linfeng Tang, Yeda Wang, Zhanchuan Cai, Junjun Jiang, Jiayi Ma
摘要
Current image fusion methods struggle with real-world composite degradations and lack the flexibility to accommodate user-specific needs. To address this, we propose ControlFusion, a controllable fusion network guided by language-vision prompts that adaptively mitigates composite degradations. On the one hand, we construct a degraded imaging model based on physical mechanisms, such as the Retinex theory and atmospheric scattering principle, to simulate composite degradations and provide a data foundation for addressing realistic degradations. On the other hand, we devise a prompt-modulated restoration and fusion network that dynamically enhances features according to degradation prompts, enabling adaptability to varying degradation levels. To support user-specific preferences in visual quality, a text encoder is incorporated to embed user-defined degradation types and levels as degradation prompts. Moreover, a spatial-frequency collaborative visual adapter is designed to autonomously perceive degradations from source images, thereby reducing complete reliance on user instructions. Extensive experiments demonstrate that Control-Fusion outperforms SOTA fusion methods in fusion quality and degradation handling, particularly under real-world and compound degradations. The source code is publicly available at https://github.com/Linfeng-Tang/ControlFusion.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- MagicFuse: Single Image Fusion for Visual and Semantic ReinforcementHao Zhang, Yanping Zha, Zizhuo Li, Meiqi Gong 等CVPR 2026 · 被引用 1 次
- ReCoFuse: Ultra-Robust Image Fusion via Restorative Multi-Modal Diffusion Reciprocal CouplingHao Zhang, Shuhan Yang, Linfeng Tang, Xunpeng Yi 等CVPR 2026
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat 等CVPR 2022 · 被引用 3,348 次
- MUSIQ: Multi-scale Image Quality TransformerJunjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar 等ICCV 2021 · 被引用 1,325 次
- Exploring CLIP for Assessing the Look and Feel of ImagesJianyi Wang, Kelvin C. K. Chan, Chen Change LoyAAAI 2023 · 被引用 1,208 次
- Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object DetectionJinyuan Liu, Xin Fan, Zhanbo Huang, Guanyao Wu 等CVPR 2022 · 被引用 929 次
相关 Paper
- Robust Fusion Controller: Degradation-Aware Image Fusion with Fine-Grained Language InstructionsHao Zhang, Yanping Zha, Qingwei Zhuang, Zhenfeng Shao 等AAAI 2026
- Text-DiFuse: An Interactive Multi-Modal Image Fusion Framework based on Text-modulated Diffusion ModelHao Zhang, Lei Cao, Jiayi MaNeurIPS 2024 · 被引用 69 次
- Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image FusionXunpeng Yi, Han Xu, Hao Zhang, Linfeng Tang 等CVPR 2024 · 被引用 121 次
- MMAIF: Multi-Task and Multi-Degradation All-in-One for Image Fusion with Language GuidanceZihan Cao, Yu Zhong, Ziqi Wang, Liang-Jian DengICCV 2025 · 被引用 1 次
- CtrlFuse: Mask-Prompt Guided Controllable Infrared and Visible Image FusionYiming Sun, Yuan Ruan, Qinghua Hu, Pengfei ZhuAAAI 2026
