Multigrain-aware Semantic Prototype Scanning and Tri-Token Prompt Learning Embraced High-Order RWKV for Pan-Sharpening
Junfeng Li, Wenyang Zhou, Xueheng Li, Xuanhua He, Jianhou Gan, Wenqi Ren
Abstract
In this work, we propose a Multigrain-aware Semantic Prototype Scanning paradigm for pan-sharpening, built upon a KV-sharing RWKV architecture for efficient global modeling, coupled with a novel tri-token prompting mechanism derived from semantic clustering to steer the fusion process adhering to the following principles: 1) Multigrain-aware Semantic Prototype Scanning. While the RWKV model offers an efficient linear alternative, its recurrent scanning mechanism often introduces positional bias and lacks semantic guidance. To address this, we introduce a semantic-driven scanning strategy. Local hashing is first employed to generate semantic prototypes via clustering, segmenting the image into coherent regions. Our scanning mechanism is then explicitly aware of multi-grain semantic structures, allowing the model to focus on contextually relevant regions during fusion, thereby enhancing spectral integrity and spatial coherence beyond sequence-agnostic approaches. 2) Tri-token Prompt Learning. The core of our framework is a tri-token prompting mechanism: (i) a globally-sourced token to encapsulate the holistic image context, (ii) cluster-derived prototype tokens to represent distinct semantic regions, and (iii) learnable token register that acts as a dynamic buffer to explicitly identify and eliminate feature noisy artifacts that commonly arise from standard global modeling. The global and prototype tokens are broadcast as semantic prompts to guide RWKV's processing, while the register continuously refines the intermediate features. 3) Invertible Q-Shift. To counteract spatial detail, we tailor two key designs: apply a center difference convolution on value pathway within the RWKV block, actively injecting high-frequency information to preserve fine textures and moving beyond parameter-heavy receptive field expansion via invertible neural network empowered multi-scale Q-shift operation. This module performs efficient, lossless feature transformation and shifting across split channels, significantly enriching feature representation. Experimental results demonstrate superiority of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0a3d4335-8d46-403c-a184-16244a536774Builds on6
- Mutual Information-driven Pan-sharpeningMan Zhou, Keyu Yan, Jie Huang, Zihe Yang et al.CVPR 2022 · 113 citations
- PanFlowNet: A Flow-Based Deep Network for Pan-sharpeningGang Yang, Xiangyong Cao, Wenzhe Xiao, Man Zhou et al.ICCV 2023 · 43 citations
- Omnidirectional Image Super-resolution via Bi-projection FusionJiangang Wang, Yuning Cui, Yawen Li, Wenqi Ren et al.AAAI 2024 · 15 citations
- Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like ArchitecturesYuchen Duan, Weiyun Wang, Zhe Chen, Xizhou Zhu et al.ICLR 2025 · 10 citations
- RAP-SR: RestorAtion Prior Enhancement in Diffusion Models for Realistic Image Super-ResolutionJiangang Wang, Qingnan Fan, Jinwei Chen, Hong Gu et al.AAAI 2025 · 4 citations
Related papers
- WKV-sharing embraced random shuffle RWKV high-order modeling for pan-sharpeningMan Zhou, Xuanhua He, Danfeng Hong, Bo HuangNeurIPS 2025 · 2 citations
- Freq-RWKV: Granularity-Aware Spatial-Frequency Synergy via Dual-Domain Recurrent Scanning for Pan-sharpeningXueheng Li, Xuanhua He, Tao Hu, Jie Zhang et al.ACM MM 2025 · 1 citation
- Hierarchical Dual-Domain Fusion with Frequency-Guided Spatial Modeling for Pan-SharpeningHuangqimei Zheng, Chengyi Pan, Qian Jiang, Wei Zhou et al.AAAI 2026
- Pan-Sharpening with Customized Transformer and Invertible Neural NetworkMan Zhou, Jie Huang, Yanchi Fang, Xueyang Fu et al.AAAI 2022 · 130 citations
- Cross-Scale Pansharpening via ScaleFormer and the PanScale BenchmarkKe Cao, Xuanhua He, Xueheng Li, Lingting Zhu et al.CVPR 2026 · 4 citations
