Empowering Semantic-Sensitive Underwater Image Enhancement with VLM
Guodong Fan, Shengning Zhou, Genji Yuan, Huiyu Li, Jingchun Zhou, Jinjiang Li
Abstract
In recent years, learning-based underwater image enhancement (UIE) techniques have rapidly evolved. However, distribution shifts between high-quality enhanced outputs and natural images can hinder semantic cue extraction for downstream vision tasks, thereby limiting the adaptability of existing enhancement models. To address this challenge, this work proposes a new learning mechanism that leverages Vision-Language Models (VLMs) to empower UIE models with semantic-sensitive capabilities. To be concrete, our strategy first generates textual descriptions of key objects from a degraded image via a VLM. Subsequently, a text-image alignment model remaps these relevant descriptions back onto the image to produce a spatial semantic guidance map. This map then steers the UIE network through a dual-guidance mechanism, which combines cross-attention and an explicit alignment loss. This forces the network to focus its restorative power on semantic-sensitive regions during image reconstruction, rather than pursuing a globally uniform improvement, thereby ensuring the faithful restoration of key object features. Experiments confirm that when our strategy is applied to different UIE baselines, significantly boosts their performance on perceptual quality metrics as well as enhances their performance on detection and segmentation tasks, validating its effectiveness and adaptability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9929dc16-ed67-47cc-90bf-92875b0c2c41Builds on7
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- PromptRestorer: A Prompting Image Restoration Method with Degradation PerceptionCong Wang, Jinshan Pan, Wei Wang, Jiangxin Dong et al.NeurIPS 2023 · 109 citations
- Synergistic Multiscale Detail Refinement via Intrinsic Supervision for Underwater Image EnhancementDehuan Zhang, Jingchun Zhou, Chunle Guo, Weishi Zhang et al.AAAI 2024 · 55 citations
- Gradient-based Visual Explanation for Transformer-based CLIPChenyang Zhao, Kun Wang, Xingyu Zeng, Rui Zhao et al.ICML 2024 · 24 citations
Related papers
- Conditional Prompt Learning via Degradation Perception for Underwater Image EnhancementMingze Yao, Zhiying Jiang, Xianping Fu, Huibing WangAAAI 2026
- Adaptive Dual-domain Learning for Underwater Image EnhancementLintao Peng, Liheng BianAAAI 2025 · 9 citations
- DuSSS: Dual Semantic Similarity-Supervised Vision-Language Model for Semi-Supervised Medical Image SegmentationQingtao Pan, Wenhao Qiao, Jingjiao Lou, Bing Ji et al.AAAI 2025 · 13 citations
- UPRE: Zero-Shot Domain Adaptation for Object Detection via Unified Prompt and Representation EnhancementXiao Zhang, Fei Wei, Yong Wang, Wenda Zhao et al.ICCV 2025 · 1 citation
- UARE: A Unified Vision-Language Model for Image Quality Assessment, Restoration, and EnhancementWeiqi Li, Xuanyu Zhang, Bin Chen, Jingfen Xie et al.CVPR 2026 · 5 citations
