CLIPPan: Adapting CLIP as a Supervisor for Unsupervised Pansharpening
Lihua Jian, Jiabo Liu, Shaowu Wu, Lihui Chen
Abstract
Despite remarkable advancements in supervised pansharpening neural networks, these methods face domain adaptation challenges of resolution due to the intrinsic disparity between simulated reduced-resolution training data and real-world full-resolution scenarios. To bridge this gap, we propose an unsupervised pansharpening framework, CLIPPan, that enables model training at full resolution directly by taking CLIP, a visual-language model, as a supervisor. However, directly applying CLIP to supervise pansharpening remains challenging due to its inherent bias toward natural images and limited understanding of pansharpening tasks. Therefore, we first introduce a lightweight fine-tuning pipeline that adapts CLIP to recognize low-resolution multispectral, panchromatic, and high-resolution multispectral images, as well as to understand the pansharpening process. Then, building on the adapted CLIP, we formulate a novel loss integrating semantic language constraints, which aligns image-level fusion transitions with protocol-aligned textual prompts (e.g., Wald's or Khan's descriptions), thus enabling CLIPPan to use language as a powerful supervisory signal and guide fusion learning without ground truth. Extensive experiments demonstrate that CLIPPan consistently improves spectral and spatial fidelity across various pansharpening backbones on real-world datasets, setting a new state of the art for unsupervised full-resolution pansharpening.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 62b7bd3b-7759-4d2c-922b-419fc57eab8cBuilds on4
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- GeoCLIP: Clip-Inspired Alignment between Locations and Images for Effective Worldwide Geo-localizationVicente Vivanco Cepeda, Gaurav Kumar Nayak, Mubarak ShahNeurIPS 2023 · 303 citations
- Linearly-evolved Transformer for Pan-sharpeningJunming Hou, Zihan Cao, Naishan Zheng, Xuan Li et al.ACM MM 2024 · 22 citations
- Unsupervised Pan-Sharpening via Mutually Guided Detail RestorationHuangxing Lin, Yuhang Dong, Xinghao Ding, Tianpeng Liu et al.AAAI 2024 · 9 citations
Related papers
- Unsupervised Domain Adaptation for Referring Semantic SegmentationHaonan Shi, Wenwen Pan, Zhou Zhao, Mingmin Zhang et al.ACM MM 2023 · 5 citations
- DiffCLIP: Few-shot Language-driven Multimodal ClassifierJiaqing Zhang, Mingxiang Cao, Xue Yang, Kai Jiang et al.AAAI 2025 · 4 citations
- Split to Merge: Unifying Separated Modalities for Unsupervised Domain AdaptationXinyao Li, Yuke Li, Zhekai Du, Fengling Li et al.CVPR 2024
- Anchor-based Robust Finetuning of Vision-Language ModelsJinwei Han, Zhiwen Lin, Zhongyisun Sun, Yingguo Gao et al.CVPR 2024
- ATAS: Any-to-Any Self-Distillation for Enhanced Open-Vocabulary Dense PredictionJuan Yeo, Soonwoo Cha, Jiwoo Song, Hyunbin Jin et al.ICCV 2025
