Cross-Scale Pansharpening via ScaleFormer and the PanScale Benchmark
Ke Cao, Xuanhua He, Xueheng Li, Lingting Zhu, Yingying Wang, Ao Ma, Zhanjie Zhang, Man Zhou, Chengjun Xie, Jie Zhang
Abstract
Pansharpening aims to generate high-resolution multispectral images by fusing the spatial detail of panchromatic images with the spectral richness of low-resolution MS data. However, most existing methods are evaluated under limited, low-resolution settings, limiting their generalization to realworld, high-resolution scenarios. To bridge this gap, we systematically investigate the data, algorithmic, and computational challenges of cross-scale pansharpening. We first introduce PanScale, the first large-scale, cross-scale pansharpening dataset, accompanied by PanScale-Bench, a comprehensive benchmark for evaluating generalization across varying resolutions and scales. To realize scale generalization, we propose ScaleFormer, a novel architecture designed for multi-scale pansharpening. ScaleFormer reframes generalization across image resolutions as generalization across sequence lengths: it tokenizes images into patch sequences of the same resolution but variable length proportional to image scale. A Scale-Aware Patchify module enables training for such variations from fixed-size crops. ScaleFormer then decouples intra-patch spatial feature learning from inter-patch sequential dependency modeling, incorporating Rotary Positional Encoding to enhance extrapolation to unseen scales. Extensive experiments show that our approach outperforms SOTA methods in fusion quality and cross-scale generalization. The datasets and source code are available at https://github.com/caoke-963/ScaleFormer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dfdbef73-f563-4cce-a0dd-2203d130ab79Builds on14
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image SynthesisJunsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao et al.ICLR 2024 · 831 citations
- Pan-Sharpening with Customized Transformer and Invertible Neural NetworkMan Zhou, Jie Huang, Yanchi Fang, Xueyang Fu et al.AAAI 2022 · 130 citations
- Mutual Information-driven Pan-sharpeningMan Zhou, Keyu Yan, Jie Huang, Zihe Yang et al.CVPR 2022 · 113 citations
- PanFlowNet: A Flow-Based Deep Network for Pan-sharpeningGang Yang, Xiangyong Cao, Wenzhe Xiao, Man Zhou et al.ICCV 2023 · 43 citations
Related papers
- HyperTransformer: A Textural and Spectral Feature Fusion Transformer for PansharpeningWele Gedara Chaminda Bandara, Vishal M. PatelCVPR 2022 · 175 citations
- Hierarchical Dual-Domain Fusion with Frequency-Guided Spatial Modeling for Pan-SharpeningHuangqimei Zheng, Chengyi Pan, Qian Jiang, Wei Zhou et al.AAAI 2026
- CTCP: Cross Transformer and CNN for PansharpeningZhao Su, Yong Yang, Shuying Huang, Weiguo Wan et al.ACM MM 2023 · 7 citations
- Adaptively Learning Low-high Frequency Information Integration for Pan-sharpeningMan Zhou, Jie Huang, Chongyi Li, Hu Yu et al.ACM MM 2022 · 44 citations
- Multi-scale Spatial-Spectral Attention Guided Fusion Network for PansharpeningYong Yang, Mengzhen Li, Shuying Huang, Hangyuan Lu et al.ACM MM 2023 · 19 citations
