DIFFSSR: Stereo Image Super-resolution Using Differential Transformer
Dafeng Zhang
Abstract
In the field of computer vision, the task of stereo image super-resolution (StereoSR) has garnered significant attention due to its potential applications in augmented reality, virtual reality, and autonomous driving. Traditional Transformer-based models, while powerful, often suffer from attention noise, leading to suboptimal reconstruction issues in super-resolved images. This paper introduces DIFFSSR, a novel neural network architecture designed to address these challenges. We introduce the Diff Cross Attention Block (DCAB) and the Sliding Stereo Cross-Attention Module (SSCAM) to enhance feature integration and mitigate the impact of attention noise. The DCAB differentiates between relevant and irrelevant context, amplifying attention to important features and canceling out noise. The SSCAM, with its sliding window mechanism and disparity-based attention, adapts to local variations in stereo images, preserving details, and addressing the performance degradation due to misalignment of horizontal epipolar lines in stereo images. Extensive experiments on benchmark datasets demonstrate that DIFF-SSR outperforms state-of-the-art methods, including NAFSSR and SwinFIRSSR, in terms of both quantitative metrics and visual quality. Code is available at https://github.com/Zdafeng/DIFFSSR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Fast Fourier ConvolutionLu Chi, Borui Jiang, Yadong MuNeurIPS 2020 · 842 citations
- Feedback Network for Mutually Boosted Stereo Image Super-Resolution and Disparity EstimationQinyan Dai, Juncheng Li, Qiaosi Yi, Faming Fang et al.ACM MM 2021 · 68 citations
- SIR-Former: Stereo Image Restoration Using TransformerZizheng Yang, Mingde Yao, Jie Huang, Man Zhou et al.ACM MM 2022 · 24 citations
- Differential TransformerTianzhu Ye, Li Dong, Yuqing Xia, Yutao Sun et al.ICLR 2025
Related papers
- DynamicStereo: Consistent Dynamic Depth from Stereo VideosNikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova et al.CVPR 2023
- Stereo Video Super-Resolution via Exploiting View-Temporal CorrelationsRuikang Xu, Zeyu Xiao, Mingde Yao, Yueyi Zhang et al.ACM MM 2021 · 20 citations
- StereoINR: Cross-View Geometry Consistent Stereo Super Resolution with Implicit Neural RepresentationYi Liu, Xinyi Liu, Yi Wan, Panwang Xia et al.ACM MM 2025
- Learning Texture Transformer Network for Image Super-ResolutionFuzhi Yang, Huan Yang, Jianlong Fu, Hongtao Lu et al.CVPR 2020
- TransMVSNet: Global Context-aware Multi-view Stereo Network with TransformersYikang Ding, Wentao Yuan, Qingtian Zhu, Haotian Zhang et al.CVPR 2022 · 236 citations
