Dual-Window Multiscale Transformer for Hyperspectral Snapshot Compressive Imaging
Fulin Luo, Xi Chen, Xiuwen Gong, Weiwen Wu, Tan Guo
Abstract
Coded aperture snapshot spectral imaging (CASSI) system is an effective manner for hyperspectral snapshot compressive imaging. The core issue of CASSI is to solve the inverse problem for the reconstruction of hyperspectral image (HSI). In recent years, Transformer-based methods achieve promising performance in HSI reconstruction. However, capturing both long-range dependencies and local information while ensuring reasonable computational costs remains a challenging problem. In this paper, we propose a Transformer-based HSI reconstruction method called dual-window multiscale Transformer (DWMT), which is a coarse-to-fine process, reconstructing the global properties of HSI with the longrange dependencies. In our method, we propose a novel U-Net architecture using a dual-branch encoder to refine pixel information and full-scale skip connections to fuse different features, enhancing the extraction of fine-grained features. Meanwhile, we design a novel self-attention mechanism called dual-window multiscale multi-head self-attention (DWM-MSA), which utilizes two different-sized windows to compute self-attention, which can capture the long-range dependencies in a local region at different scales to improve the reconstruction performance. We also propose a novel position embedding method for Transformer, named con-abs position embedding (CAPE), which effectively enhances positional information of the HSIs. Extensive experiments on both the simulated and the real data are conducted to demonstrate the superior performance, stability, and generalization ability of our DWMT. Code of this project is at https://github.com/chenx2000/DWMT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d6bc395d-9e80-4074-bc10-eb60f93b5b5fCited by top-tier papers2
- Detail Matters: Mamba-Inspired Joint Unfolding Network for Snapshot Spectral Compressive ImagingMengjie Qin, Yuchao Feng, Zongliang Wu, Yulun Zhang et al.AAAI 2025 · 21 citations
- Bridging Local–Global Dissonance: Learning from Compressive Measurements for Hyperspectral ReconstructionXian-Hua HanICML 2026
Builds on14
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Vision Transformer with Deformable AttentionZhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li et al.CVPR 2022 · 835 citations
- Understanding Robustness of Transformers for Image ClassificationSrinadh Bhojanapalli, Ayan Chakrabarti, Daniel Glasner, Daliang Li et al.ICCV 2021 · 501 citations
- Fast Convergence of DETR with Spatially Modulated Co-AttentionPeng Gao, Minghang Zheng, Xiaogang Wang, Jifeng Dai et al.ICCV 2021 · 392 citations
Related papers
- Mask-guided Spectral-wise Transformer for Efficient Hyperspectral Image ReconstructionYuanhao Cai, Jing Lin, Xiaowan Hu, Haoqian Wang et al.CVPR 2022 · 310 citations
- Degradation-Aware Unfolding Half-Shuffle Transformer for Spectral Compressive ImagingYuanhao Cai, Jing Lin, Haoqian Wang, Xin Yuan et al.NeurIPS 2022 · 222 citations
- Spectral Compressive Imaging via Chromaticity-Intensity DecompositionXiaodong Wang, Zijun He, Ping Wang, Lishun Wang et al.NeurIPS 2025 · 4 citations
- VmambaSCI: Dynamic Deep Unfolding Network with Mamba for Compressive Spectral ImagingMingjin Zhang, Longyi Li, Wenxuan Shi, Jie Guo et al.ACM MM 2024 · 13 citations
- Joint Spectral Image Reconstruction and Semantic Segmentation with Cooperative UnfoldingZijun He, Ping Wang, Xiaodong Wang, Chang Chen et al.CVPR 2026
