Multigrain-aware Semantic Prototype Scanning and Tri-Token Prompt Learning Embraced High-Order RWKV for Pan-Sharpening
Junfeng Li, Wenyang Zhou, Xueheng Li, Xuanhua He, Jianhou Gan, Wenqi Ren
摘要
In this work, we propose a Multigrain-aware Semantic Prototype Scanning paradigm for pan-sharpening, built upon a KV-sharing RWKV architecture for efficient global modeling, coupled with a novel tri-token prompting mechanism derived from semantic clustering to steer the fusion process adhering to the following principles: 1) Multigrain-aware Semantic Prototype Scanning. While the RWKV model offers an efficient linear alternative, its recurrent scanning mechanism often introduces positional bias and lacks semantic guidance. To address this, we introduce a semantic-driven scanning strategy. Local hashing is first employed to generate semantic prototypes via clustering, segmenting the image into coherent regions. Our scanning mechanism is then explicitly aware of multi-grain semantic structures, allowing the model to focus on contextually relevant regions during fusion, thereby enhancing spectral integrity and spatial coherence beyond sequence-agnostic approaches. 2) Tri-token Prompt Learning. The core of our framework is a tri-token prompting mechanism: (i) a globally-sourced token to encapsulate the holistic image context, (ii) cluster-derived prototype tokens to represent distinct semantic regions, and (iii) learnable token register that acts as a dynamic buffer to explicitly identify and eliminate feature noisy artifacts that commonly arise from standard global modeling. The global and prototype tokens are broadcast as semantic prompts to guide RWKV's processing, while the register continuously refines the intermediate features. 3) Invertible Q-Shift. To counteract spatial detail, we tailor two key designs: apply a center difference convolution on value pathway within the RWKV block, actively injecting high-frequency information to preserve fine textures and moving beyond parameter-heavy receptive field expansion via invertible neural network empowered multi-scale Q-shift operation. This module performs efficient, lossless feature transformation and shifting across split channels, significantly enriching feature representation. Experimental results demonstrate superiority of our method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Mutual Information-driven Pan-sharpeningMan Zhou, Keyu Yan, Jie Huang, Zihe Yang 等CVPR 2022 · 被引用 113 次
- PanFlowNet: A Flow-Based Deep Network for Pan-sharpeningGang Yang, Xiangyong Cao, Wenzhe Xiao, Man Zhou 等ICCV 2023 · 被引用 43 次
- Omnidirectional Image Super-resolution via Bi-projection FusionJiangang Wang, Yuning Cui, Yawen Li, Wenqi Ren 等AAAI 2024 · 被引用 15 次
- Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like ArchitecturesYuchen Duan, Weiyun Wang, Zhe Chen, Xizhou Zhu 等ICLR 2025 · 被引用 10 次
- RAP-SR: RestorAtion Prior Enhancement in Diffusion Models for Realistic Image Super-ResolutionJiangang Wang, Qingnan Fan, Jinwei Chen, Hong Gu 等AAAI 2025 · 被引用 4 次
相关 Paper
- WKV-sharing embraced random shuffle RWKV high-order modeling for pan-sharpeningMan Zhou, Xuanhua He, Danfeng Hong, Bo HuangNeurIPS 2025 · 被引用 2 次
- Freq-RWKV: Granularity-Aware Spatial-Frequency Synergy via Dual-Domain Recurrent Scanning for Pan-sharpeningXueheng Li, Xuanhua He, Tao Hu, Jie Zhang 等ACM MM 2025 · 被引用 1 次
- Hierarchical Dual-Domain Fusion with Frequency-Guided Spatial Modeling for Pan-SharpeningHuangqimei Zheng, Chengyi Pan, Qian Jiang, Wei Zhou 等AAAI 2026
- Pan-Sharpening with Customized Transformer and Invertible Neural NetworkMan Zhou, Jie Huang, Yanchi Fang, Xueyang Fu 等AAAI 2022 · 被引用 130 次
- Cross-Scale Pansharpening via ScaleFormer and the PanScale BenchmarkKe Cao, Xuanhua He, Xueheng Li, Lingting Zhu 等CVPR 2026 · 被引用 4 次
