PAN-Crafter: Learning Modality-Consistent Alignment for Pan-Sharpening
Jeonghyeok Do, Sungpyo Kim, Geunhyuk Youk, Jaehyup Lee, Munchurl Kim
Abstract
PAN-sharpening aims to fuse high-resolution panchromatic (PAN) images with low-resolution multi-spectral (MS) images to generate high-resolution multi-spectral (HRMS) outputs. However, cross-modality misalignment -- caused by sensor placement, acquisition timing, and resolution disparity -- induces a fundamental challenge. Conventional deep learning methods assume perfect pixel-wise alignment and rely on per-pixel reconstruction losses, leading to spectral distortion, double edges, and blurring when misalignment is present. To address this, we propose PAN-Crafter, a modality-consistent alignment framework that explicitly mitigates the misalignment gap between PAN and MS modalities. At its core, Modality-Adaptive Reconstruction (MARs) enables a single network to jointly reconstruct HRMS and PAN images, leveraging PAN's high-frequency details as auxiliary self-supervision. Additionally, we introduce Cross-Modality Alignment-Aware Attention (CM3A), a novel mechanism that bidirectionally aligns MS texture to PAN structure and vice versa, enabling adaptive feature refinement across modalities. Extensive experiments on multiple benchmark datasets demonstrate that our PAN-Crafter outperforms the most recent state-of-the-art method in all metrics, even with 50.11 faster inference time and 0.63 the memory size. Furthermore, it demonstrates strong generalization performance on unseen satellite datasets, showing its robustness across different conditions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 05644e8d-8403-4c7a-b936-4a4ff564b9bfBuilds on21
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- HyperTransformer: A Textural and Spectral Feature Fusion Transformer for PansharpeningWele Gedara Chaminda Bandara, Vishal M. PatelCVPR 2022 · 175 citations
- LAGConv: Local-Context Adaptive Convolution Kernels with Global Harmonic Bias for PansharpeningZi-Rong Jin, Tian-Jing Zhang, Tai-Xiang Jiang, Gemine Vivone et al.AAAI 2022 · 131 citations
Related papers
- SIPSA-Net: Shift-Invariant Pan Sharpening With Moving Object Alignment for Satellite ImageryJaehyup Lee, Soomin Seo, Munchurl KimCVPR 2021
- Adaptively Learning Low-high Frequency Information Integration for Pan-sharpeningMan Zhou, Jie Huang, Chongyi Li, Hu Yu et al.ACM MM 2022 · 44 citations
- Multi-scale Spatial-Spectral Attention Guided Fusion Network for PansharpeningYong Yang, Mengzhen Li, Shuying Huang, Hangyuan Lu et al.ACM MM 2023 · 19 citations
- Learning High-frequency Feature Enhancement and Alignment for Pan-sharpeningYingying Wang, Yunlong Lin, Ge Meng, Zhenqi Fu et al.ACM MM 2023 · 21 citations
- Hierarchical Dual-Domain Fusion with Frequency-Guided Spatial Modeling for Pan-SharpeningHuangqimei Zheng, Chengyi Pan, Qian Jiang, Wei Zhou et al.AAAI 2026
