Xformer: Hybrid X-Shaped Transformer for Image Denoising
Jiale Zhang, Yulun Zhang, Jinjin Gu, Jiahua Dong, Linghe Kong, Xiaokang Yang
摘要
In this paper, we present a hybrid X-shaped vision Transformer, named Xformer, which performs notably on image denoising tasks. We explore strengthening the global representation of tokens from different scopes. In detail, we adopt two types of Transformer blocks. The spatial-wise Transformer block performs fine-grained local patches interactions across tokens defined by spatial dimension. The channel-wise Transformer block performs direct global context interactions across tokens defined by channel dimension. Based on the concurrent network structure, we design two branches to conduct these two interaction fashions. Within each branch, we employ an encoder-decoder architecture to capture multi-scale features. Besides, we propose the Bidirectional Connection Unit (BCU) to couple the learned representations from these two branches while providing enhanced information fusion. The joint designs make our Xformer powerful to conduct global information modeling in both spatial and channel dimensions. Extensive experiments show that Xformer, under the comparable model complexity, achieves state-of-the-art performance on the synthetic and realworld image denoising tasks. We also provide code and models at https: //github.com/gladzhang/Xformer .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- DreamClear: High-Capacity Real-World Image Restoration with Privacy-Safe Dataset CurationYuang Ai, Xiaoqiang Zhou, Huaibo Huang, Xiaotian Han 等NeurIPS 2024 · 被引用 81 次
- Sharing Key Semantics in Transformer Makes Efficient Image RestorationBin Ren, Yawei Li, Jingyun Liang, Rakesh Ranjan 等NeurIPS 2024 · 被引用 17 次
- Beyond the Ground Truth: Enhanced Supervision for Image RestorationDonghun Ryou, Inju Ha, Sanghyeok Chu, Bohyung HanCVPR 2026 · 被引用 4 次
- StreamFlow: Theory, Algorithm, and Implementation for High-Efficiency Rectified Flow GenerationSen Fang, Hongbin Zhong, Yalin Feng, Yanxin Zhang 等ICML 2026 · 被引用 3 次
- Self-Calibrated Variance-Stabilizing Transformations for Real-World Image DenoisingSébastien Herbreteau, Michael UnserICCV 2025 · 被引用 3 次
它引用的顶会 Paper24
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat 等CVPR 2022 · 被引用 3,348 次
相关 Paper
- Uformer: A General U-Shaped Transformer for Image RestorationZhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou 等CVPR 2022 · 被引用 1,970 次
- Activating More Pixels in Image Super-Resolution TransformerXiangyu Chen, Xintao Wang, Jiantao Zhou, Yu Qiao 等CVPR 2023
- Conformer: Local Features Coupling Global Representations for Visual RecognitionZhiliang Peng, Wei Huang, Shanzhi Gu, Lingxi Xie 等ICCV 2021 · 被引用 723 次
- Hybrid CNN-Transformer Feature Fusion for Single Image DerainingXiang Chen, Jinshan Pan, Jiyang Lu, Zhentao Fan 等AAAI 2023 · 被引用 75 次
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 被引用 2,072 次
